BaZi · Blog

BaZi Celebrity Census: Sample Limits

Deep Oracle Editorial Team · 2026-08-13 · 8 min read

BaZi Celebrity Census: Sample Limits: Corpus Size

The current BaZi celebrity repository contains 2,223 JSON files on disk as of 2026-08-21. Of these, 2,017 profiles pass the live loader's chart-integrity publish gate. Among publishable profiles, only 47 have a known hour, while 1,970 have an unknown hour. Additionally, 1,524 publishable profiles carry a Wikidata QID and at least one Wikidata reference. This corpus is a curated public-profile dataset, not a random or representative population sample. These counts reflect dataset scope and quality controls rather than prevalence or causal patterns.

Publish-Gated Sample

Each profile must pass the live loader’s chart-integrity publish gate. This gate verifies field-by-field: sex, year pillar, month pillar, day pillar, hour pillar (if claimed), and stem-branch consistency. A known-hour profile is rejected if the hour pillar conflicts with the day stem; an unknown-hour profile must omit the hour pillar entirely or it fails. The gate also requires a Wikidata QID and at least one Wikidata reference for every publishable profile. On 2026-08-21, 2,017 of 2,223 celebrity JSON files passed; 47 had a known hour, 1,970 had an unknown hour, and 1,524 carried both a QID and a reference. Profiles failing any check remain on disk but are excluded from publication. This is a curated public-profile dataset, not a random or representative population sample. See relevant page.

BaZi Celebrity Census: Sample Limits: Birth-Hour Coverage Verification

When assessing birth-hour coverage, verify each profile against the live loader’s chart-integrity publish gate rather than the raw file count. On 2026-08-21, the repository census shows 2,223 celebrity JSON files on disk, but only 2,017 pass the publish gate. Of these publishable profiles, 47 have a known hour while 1,970 have an unknown hour. This means a chart may exist as a file yet lack sufficient hour data to be published. Always confirm the hour field is present and non-null; otherwise, treat the profile as unpublishable for hour-based analysis. The corpus is a curated public-profile dataset, not a random or representative population sample, so hour coverage cannot be generalized. For guidance on interpreting charts with missing hour, see relevant page.

Source-Field Coverage

Every profile in the corpus is checked against its source fields before it can pass the loader’s publish gate. The loader reads the raw JSON for the name, birth date, and, when present, the hour field; it then compares these against the chart’s own generated pillars. A profile with a known hour must align that hour with the day stem and branch, while unknown-hour profiles are verified against the day and month only. Because the dataset contains far more unknown-hour records (1,970) than known-hour ones (47), the loader treats hour as a boundary case: it never imputes a missing hour, and it rejects any file where the hour conflicts with the calculated day pillar. This field-by-field check is what keeps the publishable count distinct from the raw file count. For more on how day strength is assessed, see relevant page.

Sample-versus-population boundary check

A field-by-field check protects the boundary between this curated sample and any population claim. For each record, confirm chart integrity first: the 2,017 profiles that pass the live loader’s publish gate are the only usable units. Then separate hour-known (47) from hour-unknown (1,970) cases, because the unknown-hour subset cannot support pillar-level statements. Finally, verify source anchoring: only the 1,524 profiles with a Wikidata QID and at least one Wikidata reference may carry a citable claim. The remaining 2,223 files on disk include records that fail one of these gates, so they are excluded from methodology totals. This procedure ensures the corpus is treated as a public-profile dataset, never as a random or representative sample. For a full discussion of how pillar timing affects analysis, see relevant page.

To verify which questions the census can answer, begin by confirming the loader gate for every profile: a chart is counted only if it passes chart-integrity checks, not merely by file presence. Next, record whether the hour is known or unknown, because that field changes the possible analyses. Then, for profiles with a Wikidata QID and at least one Wikidata reference, cross-check the identifier and citation independently; absence of either means the profile cannot support source-tracing questions. Finally, treat the corpus as a curated public-profile dataset, not a random or representative population sample. Any question about prevalence, causation, or life outcomes is outside scope. For field-level interpretation, consult relevant page.

Questions This Dataset Cannot Answer

This corpus of 2,223 celebrity JSON files, with 2,017 passing the live loader's chart-integrity publish gate, is a curated public-profile collection, not a random or representative population sample. Accordingly, it cannot support claims about prevalence, causation, or life outcomes. Verification of each profile requires checking that its Si Zhu pillars match the recorded birth data, that the hour branch is marked unknown when absent, and that at least one Wikidata reference supports the stated birth date. The 47 profiles with a known hour and the 1,970 without cannot together answer questions about hour distribution or demographic patterns. Any statement about general population frequencies or predictive validity lies outside this dataset's scope.

To replicate the corpus, an auditor checks each of the 2,223 celebrity JSON files on disk against the live loader’s chart-integrity publish gate; only 2,017 profiles pass. For the 47 publishable profiles with a known hour, the hour field must be non-empty and parseable; for the 1,970 with unknown hour, the field is explicitly null. Each publishable profile is inspected for a four-pillar structure derived from a birth date; malformed or missing pillars fail the gate. Boundary cases include files present on disk but excluded by the loader, leaving 206 non-publishable records. For Wikidata alignment, 1,524 publishable profiles require at least one reference and a QID; the remaining publishable profiles are noted as lacking these identifiers. The procedure does not infer representativeness or prevalence from any count.

Counting Definition

Each published record is counted only if it survives the loader's chart-integrity gate. We verify the four pillars, stem-branch combinations, and sex of the subject as required fields; missing values discard the record. Known hour is defined as an explicit birth-hour value in the source file, while unknown hour includes both absent and placeholder entries. A Wikidata QID is accepted when the profile contains a non-empty QID and at least one associated reference. Duplicate QIDs or repeated chart signatures within the same file are treated as one record. This field-by-field procedure bounds the published corpus: 2,223 files on disk reduce to 2,017 publishable profiles, of which 47 have a known hour and 1,524 carry a valid QID with a reference.

Data-quality limits. Each published chart is verified field by field against its source record. The loader checks the four pillars, stem-branch pairings, and any recorded hour; a profile passes the publish gate only when these values are internally consistent and match the stored JSON. Boundaries arise for unknown-hour entries: the hour pillar is omitted rather than guessed, which means 1,970 of the 2,017 publishable profiles lack that field. When a Wikidata QID is present—true for 1,524 profiles—the reference link is checked for resolvability, but the text itself is not re-derived. All counts come from the live repository census of 2026-08-21 and describe the curated celebrity corpus only; they do not support population inference or outcome claims.

BaZi Celebrity Census: Sample Limits: Reader Fact Check

To verify a published profile field by field, start with the live loader’s chart-integrity publish gate. Confirm the profile appears in the repository census: 2,223 celebrity JSON files exist on disk, while 2,017 pass the gate. Check whether the profile lists a known hour (47 profiles) or an unknown hour (1,970 profiles). For source validation, look for a Wikidata QID and at least one Wikidata reference; 1,524 publishable profiles carry both. Because the corpus is a curated public-profile dataset, not a random or representative population sample, treat any prevalence, causation, or life-outcome claim as out of scope. Boundary case: a profile with a QID but no reference fails the reference check; a profile missing a QID is not counted in the 1,524.

The study's conclusion limits are set by a field-by-field verification procedure. Each chart must pass a publish gate checking chart integrity, which yields 2,017 publishable profiles from 2,223 raw files. For every field, the loader verifies data presence and format; a missing or malformed hour field moves the profile to the unknown-hour set (1,970 profiles), while only 47 have a known hour. Verification also requires a Wikidata QID and at least one reference for 1,524 profiles; profiles without both remain outside that subset. Because the corpus is a curated public-profile dataset, not a random or representative population sample, any conclusion is bounded to data-quality statements. No prevalence, causation, or life-outcome claim can be supported. This boundary case ensures that each verified field contributes only to methodological transparency, never to inferential generalization.

The corpus size must be verified field by field, because every figure is a boundary condition, not a sample statistic. Begin with the total file count: 2,223 celebrity JSON files exist on disk. Next, apply the loader’s publish gate: only 2,017 profiles pass chart-integrity checks; the remaining 206 are excluded from any publishable set. Within those 2,017, the hour field is the primary boundary case: 47 profiles have a known hour, while 1,970 have an unknown hour, meaning hour-based analyses are constrained to a very small subset. Finally, the reference field adds a further boundary: 1,524 publishable profiles carry a Wikidata QID and at least one Wikidata reference, leaving 493 without that linkage. Each field therefore defines a distinct subcorpus. Verification must report each count separately; no aggregate can be treated as a population estimate.

BaZi Celebrity Census: Sample Limits