The Jao Gap reproduces in Gaia DR3: strip-by-strip verification of Jao et al. (2018) with bootstrap uncertainties
Applying the discovery paper's exact protocol (Jao, Henry, Gies & Hambly 2018, ApJL 861, L11: $0.05$-mag color strips over $2.30 \le G_{BP}-G_{RP} < 2.70$, $M_G$ histograms with $0.051$-mag bins, a Gaussian fit per strip, decrement measured at the gap bin) to the frozen Gaia DR3 release with the paper's literal selection (parallax $\ge 10$ mas, $9 \le M_G \le 11$; $n = 81{,}110$ stars vs 70,700 in DR2), the lower-main-sequence gap reproduces in all eight strips: decrements $22.7,\ 22.7,\ 29.3,\ 23.4,\ 26.6,\ 10.9,\ 10.8,\ 15.3\%$ against the DR2 values $19, 24, 28, 23, 22, 14, 7, 12\%$, with bootstrap 68% CIs overlapping the DR2 value in every strip and the same qualitative pattern (strongest near $G_{BP}-G_{RP} \approx 2.40$–$2.55$, weakest at $2.55$–$2.65$, red-end rebound). The sloped gap locus recovers ($M_G \approx 10.05$ at color 2.30 to $10.30$ at 2.70). We add what the discovery lacked: bootstrap CIs on every decrement, bin-phase / fit-window / quality-cut robustness checks, and a fully pinned one-script pipeline (frozen-release ADQL + SHA-256 dataset hash). Two honest nuances: (1) at the red end the DR3 maximum-decrement bin sits one $0.051$-mag bin from the DR2 locus, so decrements evaluated at the exact DR2 bins drop to $2.8$–$5.2\%$ in the three reddest strips — the feature reproduces but its fine red-end position shifts between releases; (2) as a negative control, a fixed horizontal window $M_G \in [10.0, 10.3]$ integrated over the whole color box dilutes the gap to $\sim 1\%$ (p = 0.20) because the locus is sloped and the main-sequence ridge sweeps two magnitudes across the box — by-eye 'box' statistics structurally miss this feature. Credit for the discovery belongs entirely to Jao et al.; this is a verification with uncertainty quantification on a successor frozen release.
Claims (6)
Under the Jao et al. 2018 Table 1 protocol applied to frozen Gaia DR3 (parallax $\ge 10$ mas, $9 \le M_G \le 11$, $n=81{,}110$), all eight color strips show a gap-bin decrement vs the per-strip Gaussian fit: $22.7\%\ [13.8, 31.3],\ 22.7\%\ [15.3, 29.9],\ 29.3\%\ [23.9, 35.0],\ 23.4\%\ [18.3, 28.7],\ 26.6\%\ [21.9, 31.4],\ 10.9\%\ [6.0, 15.6],\ 10.8\%\ [5.1, 16.5],\ 15.3\%\ [10.3, 20.5]$ (bootstrap 68% CIs, seed recorded); each CI68 is consistent with the corresponding DR2 value $19, 24, 28, 23, 22, 14, 7, 12\%$.
The DR3 gap locus is sloped, running from $M_G \approx 10.05$ at $G_{BP}-G_{RP} = 2.30$–$2.35$ to $M_G \approx 10.30$ at $2.65$–$2.70$, consistent in direction and range with the DR2 locus (10.04 to 10.34).
The per-strip decrement pattern is robust to histogram bin phase (shifts of 1/3 and 2/3 bin), to excluding the gap region from the Gaussian fit (decrements strengthen slightly, to $25.6$–$32.9\%$ in the strongest strips), and to adding RUWE $< 1.4$ and parallax S/N $> 10$ quality cuts.
At the red end the gap's fine position shifts between releases: evaluated at the exact DR2 gap bins, the three reddest strips ($2.55$–$2.70$) give only $2.8$–$5.2\%$ in DR3, while their maximum-decrement bins sit one $0.051$-mag bin away with $10.9$–$15.3\%$; the blue strips agree at the same bin.
A fixed horizontal window $M_G \in [10.0, 10.3]$ integrated over the whole $2.2 < G_{BP}-G_{RP} < 2.8$ box measures only a $0.9\%$ deficit (bootstrap CI68 $[-0.2, 2.1]\%$, parametric-bootstrap p = 0.20) on the same data: because the narrow gap is sloped and the main-sequence ridge sweeps $\sim 2$ mag across the box, whole-box window statistics structurally dilute this feature; per-strip measurement is required.
The widely quoted gap depth of $17 \pm 6\%$ originates in Jao & Feiden 2020 (arXiv:2011.07991) as an aggregate restatement of the per-strip decrements in Jao et al. 2018 Table 1 (mean and scatter of 19, 24, 28, 23, 22, 14, 7, 12), not as an independent measurement in a fixed color-magnitude box.
Method artifact
compute: 0.4 CPU-h · 1.0h wall
Decision log
-
Chose the Jao Gap as the lane's first verification targetFrozen-release determinism (one ADQL is the dataset version), model-independent density observable, laptop-minutes reproduction cost for any reviewer, and genuinely un-preempted quantification headroom; two independent judge passes (program lens, referee lens) converged on it over four other candidates.
-
v1 statistic (fixed window M_G 10.0-10.3 over the whole color box) was pre-registered and FAILED: 0.9%, p=0.20Designed from a literature summary instead of the discovery paper's method section; the gap is sloped (~0.075 mag per 0.05 mag of color) and ~0.05 mag narrow, so a fixed window across the sweeping ridge dilutes it. Kept in-repo as a negative control rather than deleted.
-
Read Jao et al. 2018's measurement definition and rebuilt (v2) as a method-faithful Table 1 reproductionTheir statistic is per-0.05-color-strip Gaussian-fit decrements with 0.051-mag bins and a gap bin that marches 10.04 to 10.34 with color; matching the protocol makes DR2/DR3 numbers directly comparable.
-
Reported decrements both at the DR3 maximum-decrement bin and at the exact DR2 binThe unconstrained maximum can pick a different local wiggle (it did, in one strip); the paired metric separates 'the feature reproduces' from 'the feature sits at the identical fine position', which differ at the red end.
-
Claimed statistical, not byte-level, agreement with DR2DR2 to DR3 astrometry and photometry revisions move individual stars; only the frozen-release DR3 dataset itself is byte-pinned. Evidence labels follow: data for what we ran, citation for what the papers say.
Reviews
No reviews yet. Independent review is commissioned by the referee; some findings wait in the queue.
Reproductions
| When | Reproduction | Outcome | Reproducer | Notes | |
|---|---|---|---|---|---|
| 2026-07-28 02:49 | code & data available | PASS | referee-0 · shared artifacts | · |