BaZi · Blog

AI Fortune-Telling Accuracy: How to Interpret and Evaluate

Deep Oracle Practitioner Desk · 2026-09-19 · 7 min read

An accuracy figure for AI-based divination, or AI算命准确率, is meaningless on its own. Before accepting any percentage, a reader should require the producer of that number to answer four questions: what layer of the system is being measured, how the sample was selected, who judged correctness, and how many cases were tested. Without those answers, the number is only a claim.

看到一个准确率数字先问四件事

Any stated accuracy rate should be able to answer four things: which layer is being tested, how the sample was selected, who judged right and wrong, and how many entries were tested in total. These four questions are not optional context. They are the minimum conditions for treating a number as a measurement rather than a promotional statement.

The first question matters because an AI divination tool may contain different functional layers. A chart-casting layer, which computes the Eight Characters from a birth time, has objectively verifiable outputs. An interpretive layer, which produces readings based on that chart, does not. A percentage that mixes these layers or fails to say which one was measured cannot be compared with another percentage that specifies its layer.

The second question matters because the composition of the test set changes what the score means. A set containing only common chart patterns produces a different kind of evidence from a set that also includes boundary dates, such as the transitions between solar terms or between hour branches. Scores from those two sets are not directly comparable, even if both are reported as accuracy.

The third question matters because the judge determines whether the number is objective or self-assessed. For the chart-casting layer, there is an objective answer for each case. For the interpretive layer, there is no single correct answer, so a score can only measure agreement or consistency, not correctness in the same sense.

The fourth question matters because sample size controls the precision of the figure. A ninety percent rate on ten cases supports a very different conclusion from a ninety percent rate on one thousand cases. The latter is a measurement; the former is an anecdote dressed in a percentage.

测试集是怎么选的

The way the test set is selected determines the meaning of the number. A test set built only from common configurations will tend to favor systems that handle those configurations well. A test set that includes boundary dates tests whether the system handles the harder edges of the calendar correctly. The two sets measure different things.

A score obtained only on common patterns cannot be compared with a score obtained on a set that includes boundary dates. The first says something about performance on typical cases. The second says something about performance across a wider range of difficulty. If a producer does not describe the test set, the reader cannot know which kind of evidence the number represents.

The selection method also determines whether the sample was chosen to represent real usage or chosen to make the system look good. A sample of convenient, clear-cut cases is not the same as a sample drawn from actual user queries. Without a stated selection rule, the number cannot be interpreted as representative of anything beyond the particular cases that happened to be tested.

评分标准由谁定

Who defines the scoring standard determines whether the number is objective or self-assessed. The chart-casting portion of an AI divination system has objective answers. Given a birth time, the ten Heavenly Stems and twelve Earthly Branches of the Four Pillars can be computed, and each case has one correct result. A scorer can check the output against that result.

The interpretive portion does not have that property. There is no single correct reading for a given chart. Different practitioners may give different readings of the same configuration, and none can be shown to be uniquely right in the way that a chart calculation can. A score for interpretation can therefore only measure consistency: whether the system gives the same reading for the same input, or whether its readings agree with a particular reference set of interpretations. That is not the same as measuring accuracy.

When the producer of the figure is also the judge of correctness, the number is self-assessed. That does not automatically make it false, but it does make it weaker than a number judged against an external, stated standard. A reader should ask who graded the outputs and whether that grader had an interest in the result.

样本量决定了什么

Sample size determines the precision of the figure. A ninety percent rate on ten samples and a ninety percent rate on one thousand samples are not equivalent evidence. The first has a wide range of plausible true performance; the second narrows that range considerably.

Ten cases are enough to show that a system can handle those ten cases. They are not enough to support a general claim about the system’s accuracy across the range of inputs it will meet in use. One thousand cases, if selected by a stated method, provide a stronger basis for a general statement, though the basis is still limited by the selection method and the scoring standard.

A reader should treat a number without a stated sample size as having unknown precision. The percentage may be reported to several decimal places, but if the denominator is unknown, the precision is illusory. A number is only as precise as its sample allows.

没有这四项的数字应该怎么看

An accuracy figure that lacks any one of the four items — the measured layer, the sample selection method, the scoring standard, or the sample size — should be treated as a claim, not as a measurement. It may be sincerely stated, and it may even happen to be true of some limited set of cases, but it does not provide a basis for comparison or for a decision.

A number that omits the measured layer leaves the reader unable to tell whether it refers to chart-casting, interpretation, or some mixture. A number that omits the sample selection method leaves the reader unable to tell what population the cases came from. A number that omits the scoring standard leaves the reader unable to tell whether correctness was judged objectively or by the producer. A number that omits the sample size leaves the reader unable to tell how much evidence the figure rests on.

In practice, most reported accuracy figures for AI divination omit at least one of these items. That does not mean the tools are useless. It means the numbers should not be given the weight that a measured result deserves. A percentage without these four items is closer to a slogan than to a statistic.

The site’s own practice is to state the test method and the sample size when publishing test results. That allows a reader to judge what the number covers and what it does not. A reader can then decide whether the measured layer, the selection rule, the scoring standard, and the sample size are adequate for the question the reader cares about.

Chart-casting accuracy can be measured objectively, because each case has one correct answer. That layer is amenable to verification by an independent party. Interpretation accuracy cannot be measured in the same way, because there is no unique correct answer at that layer. At most, one can measure consistency: agreement with a reference set, repeatability on identical inputs, or agreement among multiple runs.

A tool’s claimed accuracy and whether its chart-casting output can be independently rechecked are two separate pieces of information. The latter is easier for a user to verify directly. A user can take a birth time, compute the Four Pillars using a trusted reference, and compare the result with the tool’s output. That check does not require accepting anyone’s accuracy figure. It provides independent evidence about the layer that is objectively checkable.

AI Fortune-Telling Accuracy: How to Interpret and Evaluate