TESTNETGalaxy SIGNAL · Protocol Record · 2026-08-26
The full technical record. 120 of the network's first 263 mined blocks were paid for detections in the observation cadence's own reference position, not the primary target — a categorical error, corrected here in full, with every self-correction along the way left visible rather than smoothed over.
This document records an investigation carried out through extended collaboration between Galaxy SIGNAL's operator and Claude, an AI coding assistant: the operator directed the investigation and made every protocol decision; the code reading, data analysis, and script execution were performed by the assistant, working against this project's data and, where authorized, its production systems. First-person language below reflects the session record as it happened, not a claim of sole authorship by either party.
pre_rectification_chain.json — SHA-256:
79e7af107e853fc29390d3fffb2c4598165f41da097bea40fd25bb09d7526056. Check it against
whatever copy you have before trusting it; if it ever differs from this value, something
changed and the change has no explanation on this page.
Provenance: this is the network's own /chain output, captured
before the 2026-08-26 reset — not reconstructed afterward. Block 263's own
accepted_at timestamp (2026-08-25 19:41:14 UTC) is real live-mining history
baked into the file itself, a day before this correction shipped; a reconstruction built
after the fact would have no reason to carry that exact value.
This file alone is not enough to reproduce the 120/263 finding.
Each block's hit.file_pos is here, but the cadence position (status)
that decides B* is a field on the source dataset, not on the transmitted hit — it was
never part of what a block carries. Reproducing the finding requires this file's
file_pos values plus the public Breakthrough Listen L-band CSV this
network mines against (see /dataset for the current citable archive),
read at each of those row positions to check whether status starts with
A or B. That is the actual, complete instruction — not an
invitation to trust this page instead of checking it.
An independent audit of the PoUV-v2 mining gate found a categorical error, not a statistical one: 120 of the network's first 263 mined blocks were paid for detections sitting in the observation cadence's own reference position — not the primary target. The filter meant to prevent this checked for an _OFF suffix in the source filename, a convention specific to one older observation campaign, not the actual structure of the cadence (recorded in a different CSV field, status).
With all four hard-rejection filters corrected — real cadence position, coincidence with the reference position of the same cadence (tolerance derived from the pair's own drift and elapsed time), coincidence with a different target anywhere in the corpus, and the existing no-drift exclusion — only 14 of the 263 already-mined blocks would survive the corrected gate. That is not 14 candidate signals; it is 14 blocks that clear four deterministic rejection rules, nothing about signal significance.
Measured on the full corpus, independent of which blocks happened to be mined under the old gate: 0.3932% of the corpus's 28,858,709 hits are eligible (113,468 hits) — window exhaustion is not a real risk under current parameters (probability of exhausting all 2,000 window attempts: ~1.17×10-77, measured directly, not assumed).
The rest of this document is the evidence for those numbers, organized by finding rather than by the order in which the investigation actually happened. Self-corrections made along the way are kept visible where they occurred, not smoothed over — including two instances where the investigation's own first attempt at something was wrong (sections 4 and 8), and two live production incidents that came directly from correcting this gate (section 10).
Before the main finding, it matters to establish what corpus this gate operates on and against what published literature it can be checked — this is what section 4 later relies on to describe the corrected gate's result as convergent with, not merely similar to, prior published work.
The corpus-wide drift_rate cutoff is exactly ±1.999396 Hz/s — matching the ±2 Hz/s search parameter of Enriquez et al. 2017 (692 stars, L band), not the ±4 Hz/s of Price et al. 2020 ("1327-Star DR1"). Independent confirmation: a direct count of distinct sources with at least one primary-position (A*) detection in the full corpus gives 692 — an exact match to the Enriquez 2017 sample size.
The minimum SNR observed in manually-audited blocks is 20.07 — below Enriquez's own publication cutoff (SNR ≥ 25). This is consistent with the corpus being the raw intermediate output of turboSETI, not the results already filtered by the published paper.
The corrected gate and the Enriquez pipeline reach the same conclusion (zero surviving candidates after vetting) by independent methods — SNR cutoff 25 plus identification logic plus human inspection (theirs) versus four hard rejections applied to the intermediate product (this network's). This is convergence of conclusion, not replication of procedure, and should be described as such.
The status field is the real cadence structure, and it was never used. Verified on the obs_run with the largest time span (AGBT16A_999_223): the status field (A1_table, B1_table, etc.) is 100% populated (0 of 289,121 rows empty) and records the hit's actual position in the observing cadence — not a binary ON/OFF marker. The pattern: A is the same target repeated three times (A1=A2=A3); B is a different star at each position (B1≠B2≠B3), with no _OFF suffix anywhere — three genuine reference pointings, not "empty sky."
A temporal-gap heuristic alone fails on real data: the cadence for GJ412A is immediately followed by GJ412B's cadence with only a 365-second gap — indistinguishable from the ~330–380s gaps within a single cadence. Only the reset of status back to A1_table reveals the real boundary. The correct rule: a reset to A1_table is the cadence boundary, full stop — filename changes and time gaps are useful as consistency checks, never as the actual condition.
Filter B, as deployed, used the wrong signal entirely. It identified "reference" by an _OFF suffix in the filename — a convention from one older campaign, not the corpus's general structure. The robust definition is the cadence position itself: eligible only if status starts with A.
Applying that correct definition retroactively to the 263 already-mined blocks: 120/263 (45.63%) have their own paid hit sitting in a B* reference position — a detection recorded in a scan designed as a reference, not as the primary observation. This is the same categorical error already known from block 84 (which does carry an _OFF suffix), at the scale of nearly half the protocol's mined history.
This finding is categorical — it depends on no chosen parameter, tolerance, or threshold. It is the strongest single result in this audit. The corpus-wide B* fraction (by hit, not by scan) is 49.63%; the fraction among mined blocks is 45.63%. With n=263, a 4-point difference is sampling noise, not a real gap — this confirms that cadence position never entered candidate selection at all; it isn't that it entered slightly wrong.
Anticipating the obvious objection ("a B position could be a legitimate observation of its own star"): of the 116 distinct normalized sources among the 120 B* blocks, how many also have their own A* cadence somewhere else in the corpus? 105/116 (90.5%) never appear as A* anywhere in the full corpus — they are reference/control stars by construction, never a primary target. The 11 exceptions align exactly with the already-known older _OFF naming convention (HIP114430, HIP34017, HIP41307, HIP52369, HIP61094, HIP65714, HIP7748, HIP78072, HIP85365, HIP86032, HIP87579) — two independent methods of identifying the same thing agree with each other. Even in those 11 cases, the star is a primary target elsewhere, but that specific scan is still a reference pointing.
Four hard-rejection rules, corrected in full:
drift_rate == 0 (see section 6 for why this is corpus-specific, not universal physics)status starts with A), corrected as described in section 3B*) detection in the same real cadence, within a tolerance derived from the candidate's own drift rate and the real elapsed time to each referenceThe A1_table boundary has to be detected on unique scan positions, not on individual hits. A single real scan typically produces hundreds of hits (distinct frequencies, same mjd/status/source) — without deduplicating by (obs_run, mjd, status, source) first, each repeated hit of the same scan would look like the start of a new cadence, fracturing it.
A real bug was found and fixed at this exact point during the analysis. The first version of the script did not deduplicate — it counted hits, not scans. Result: 49 of the 263 blocks appeared to have no reference detection anywhere in their own cadence, all sharing the same suspicious pattern (a single-line cadence, only A1_table). Before accepting that as a real structural edge case, the cause was investigated: it was the fragmentation bug, not a genuine property of the data. The sequence, in order: a fail-closed rule (reject when no verifiable reference exists) was specified and applied first, and the 49 "no-reference" blocks were initially treated as a genuine corpus property — the first version of this document said exactly that. Only afterward, while investigating the cause of the 49 (not the number itself, which was already accepted), did the identical pattern across all of them reveal that the premise was wrong: not a cadence boundary, a grouping bug. The correction undid the premise, not just the count. After the fix: 0 of 263 blocks without a reference.
Spectral resolution (2.8355 Hz median, σ=0.0007 Hz, measured by regressing ChanIndx against Freq grouped by (FileID, CoarseChanNum) over 193,558 pairs — later confirmed against real file headers from two separate campaign eras, 2018A and 2019B, both giving 2.7939677238464355 Hz exactly, 187,500,000/2²⁶) answers "are two frequencies distinguishable?" — not the question D actually needs answered: "do two detections come from the same emitter?" A real terrestrial emitter does not reappear at the exact same frequency across observations separated by months — oscillator stability, thermal drift, satellite Doppler, and the pipeline's own peak-estimation error spread it over a window much wider than one channel. Resolution is the floor of what the threshold could be, not the right value.
Measured directly: distance (Hz) to the nearest different-source detection, restricted to three dense, recurring bands already identified among the B* blocks (~1240.03 MHz GNSS, ~1618.0–1618.5 MHz GLONASS/Iridium, ~1685.2–1685.9 MHz GOES weather satellite):
| Band | n hits | p50 | p90 | p99 | p100 |
|---|---|---|---|---|---|
| 1240 MHz (GNSS) | 4,735 | 0 Hz | 0 Hz | 0 Hz | 4,939 Hz (1 outlier) |
| 1618 MHz (GLONASS/Iridium) | 34,958 | 5 Hz | 17 Hz | 42 Hz | 536 Hz |
| 1685 MHz (GOES) | 64,212 | 0 Hz | 9 Hz | 48 Hz | 570 Hz |
p99 converges to the same order of magnitude (42–48 Hz) across the two large-sample bands; the third (1240 MHz) is the most concentrated of the three and sits well inside a 50Hz threshold. A threshold of 50 Hz, anchored to measured p99 dispersion, replaces an earlier, looser 200 Hz choice that the data don't actually support — the earlier value came from an overall distance distribution that turned out to have no natural bimodal boundary at all (a random sample of 5,000 corpus-wide hits: p0 through p75 = 0 Hz; the distribution decays continuously, with no gap separating two populations — any fixed threshold is an engineering choice over a continuous curve, not a structure discovered in the data). The p99 percentile choice (over p95 or p50) reflects a deliberate cost asymmetry: since the corpus is effectively all RFI, missing a real coincidence (too tight a threshold) is more costly than being slightly too permissive.
Scope limitation of this measurement, stated plainly: both large-sample bands measured are satellite emitters. No radar band was measured, and the 1211–1400 MHz range (which includes one of the final survivors, discussed below, classified as radar) is the third-most-represented band in the corpus. The threshold is calibrated on satellite emitters and assumed transferable to other emitter classes, without direct verification.
Verification that D's distance-zero cases are genuine physical coincidences, not artifacts. For 17 blocks with a measured distance of exactly 0 Hz, FileID and ChanIndx were recovered to distinguish three hypotheses: same cell (duplication), same FileID with a different ChanIndx (rounding), or a genuinely different FileID (real physical coincidence). Result: 100% "different FileID" — zero duplication, zero rounding artifacts. In several cases the ChanIndx is identical between completely separate observation files, once 26 days apart, once 109 days apart across two different campaigns.
Two earlier, superseded values appear in the investigation's own working notes and should never be repeated: a 3/263 figure derived from a still-arbitrary 200Hz distance threshold with B and C still using their old, wrong definitions; and a 10/263 figure with only D corrected.
This retrospective check is over the 263 already-mined blocks specifically — for each, whether its own hit would pass A/B/C/D under the corrected rules. Answering that requires reconstructing real cadence structure and nearest-neighbor distances from the surrounding data: 111 distinct obs_run values, 25,189,325 of the corpus's 28,858,709 rows (87%), scanned in two passes (count, then fill pre-allocated numpy arrays — no intermediate Python lists, a lesson from a real incident described in section 10) to compute the columns each of the 263 blocks is then looked up against. The table below is over the 263 blocks; the full-corpus-scale version of this same cascade, run independently over all 28.8M hits, is in section 7:
| Stage | Survive |
|---|---|
| initial | 263 |
| after A (drift = 0) | 183 |
| after B (position ≠ A*) | 104 |
| after C (real cadence, per-pair tolerance) | 81 |
| after D (50 Hz, exact nearest neighbor) | 14 |
Final survivors (14/263): blocks 10, 30, 35, 46, 68, 73, 104, 134, 142, 152, 153, 155, 171, 248.
Per-filter contribution in coalition (what each filter uniquely recovers if removed from the full set — robust to ordering, unlike the sequential cascade above):
| Filter removed | Survive | Uniquely recovered |
|---|---|---|
| without A | 14 | 0 |
| without B | 14 | 0 |
| without C | 17 | 3 (blocks 187, 227, 245) |
| without D | 81 | 67 |
"Without D" in this table (81) is the same number as "after C" in the sequential cascade above — not a coincidence, an identity. A, B, and C are each independent per-hit predicates, evaluated without reference to which other hits already survived, so applying A+B+C in sequence and stopping before D gives exactly the same set as applying A+B+C alone with D simply omitted. The two tables agree with each other by construction, which is itself a basic internal-consistency check on both computations.
A and B show zero unique contribution in this specific set. That is not an argument for removing them, and is stated explicitly so it is never read as one: they are definitional, not statistical — a hit at a reference position is not a primary detection regardless of any measurement, so a zero unique contribution doesn't make the rejection less correct. They are also local: C needs to reconstruct one obs_run's cadence, D needs the whole corpus to find a nearest neighbor. On the day a live data source exists, C and D have nothing to compare a brand-new hit against yet; A and B keep working exactly the same way.
D is the decisive filter — 67 of the 81 post-C survivors are rejected by D alone, which makes its threshold (above) the single most scrutinizable number in this entire document: no other measurement here decides nearly this much of the final result by itself.
Weaker than its name suggests. "ON/OFF coincidence" evokes "the same specific signal reappears in the reference scan." What C actually tests is looser: "does any detection exist in the reference within the drift-projected tolerance" — compared against every B* position in the same cadence, not a single expected counterpart. With roughly 6,200 detections per scan on average, this is considerably more permissive than classical turboSETI ON/OFF cross-checking. This explains C and D's high overlap — 20 of the 23 blocks C rejects sequentially would be rejected by D anyway, because "seen at the reference" and "coincides with another named source" end up being nearly the same question when the references are named stars, not empty sky. This is not an implementation error, but it needs to be described as the filter actually works, not as its name suggests — if C's rejection rate looks suspiciously high in a dense band after this correction ships, this is where to look first.
Unlike D (which stores one continuous per-hit invariant — distance — that any future threshold can compare against without recomputation), C's tolerance is intrinsically per-pair (depends on the candidate's own drift rate and the real elapsed time to each reference) — there is no single storable scalar for it. Changing C's formula always requires recomputing the whole c_rejected column. This is a real difference between the two filters, not an implicit symmetry.
Every node on the network independently recomputes, for every block it receives: the hit's authenticity against its own local copy of the canonical dataset, the four deterministic rejection rules above, and whether the hit falls inside the deterministic candidate window derived from the previous block's hash. All from the same public data, with no step taken on trust.
The classifier's confidence score is not on that list, on purpose. It is declared by whoever mines the block. No node — not the one validating a received block, not any other node on the network — ever independently recomputes it. The only thing checked is that the declared number falls inside a fixed range required for the block to be valid at all; within that range, the value itself is taken as given. classifier_version is in the same position: required to be present on every block, never checked against the model that actually produced the score. Both are published metadata, not verified guarantees, and this document does not suggest otherwise anywhere else in it.
Until today, a block's reward scaled with the miner's declared score (a fixed base plus a bonus proportional to it). That bonus existed to reward triage quality — a gradient. That premise doesn't hold: the four hard-rejection rules are what actually discriminates a real candidate, and they are binary — a hit clears the gate or it doesn't, nothing about passing it comes in degrees. Combined with the score never being verified, there was no real gradient left to pay for, only an unverified number every miner had every incentive to always declare at its maximum. The reward is now fixed (10 SGNL per block), regardless of the exact score attached to it. The score remains on every block as informational metadata — still required to fall in range, still with no effect on payout.
Measured consequence, not left implicit: at 113,468 eligible hits (section 7) and a fixed 10 SGNL reward, the first full pass over the corpus mints 1,134,680 SGNL. Because the reward halves every subsequent pass, the complete series (summed to infinity) converges to ~2,269,360 SGNL — about 11% of the nominal 21,000,000 SGNL supply cap — before tail emission. That cap was never a property this corpus's real eligible-hit count supports; what actually grows the supply past this point is the network's fixed perpetual tail emission, at a rate now dominant far sooner than the original design implied.
This finding only surfaced because a later, unrelated pipeline run happened to process real data from a different observation campaign (turboSETI 2.3.2 over a post-2019 cadence, outside the 2016-17 corpus this network mines) — not the kind of thing that shows up testing code against the same corpus that generated it.
Measurement: 0 hits with drift_rate == 0 in 25,321 real hits from that other campaign (production import code, unmodified) — against 56.52% of the 2016-17 corpus (16,312,152 of 28,858,709 hits).
Cause, confirmed by reading the actual turboSETI 2.3.2 source, not inferred: find_doppler.py contains, in two places (positive and negative drift search):
if abs(drift_rate) > fd.min_drift:
hitsearch(fd, spectrum, ...)
A strict comparison (>, confirmed by direct reading, not >=). With the default min_drift (0.00001, unchanged in that run), abs(0.0) > min_drift is False for any non-negative min_drift — the drift_rate = 0.0 grid point is never tested or reported, by construction, regardless of the exact min_drift chosen. This is not an artifact specific to one run; it is the behavior of that condition in any run with a non-negative min_drift.
Consequence: the older 2016-17 corpus cannot have been produced by the same code condition with a non-negative min_drift and still show 56.52% zeros — either the historical tool had different logic, or the zeros entered the published CSV by some other route. Not verifiable without the 2016-17-era source code, which isn't available — this stays unresolved, not assumed.
A real, partial explanation exists for part of the 56.52%. At least 15.8% of the total is confirmed as an instrument-processing artifact, with no natural cutoff: 509,605 hits with SNR > 10,000 concentrate on just 921 distinct frequencies (~553 hits/frequency, ~2,200 distinct sources per frequency — impossible for real sky signals, which would vary with pointing/Doppler across thousands of targets). These frequencies aren't random: they are exact dyadic subdivisions of 187.5 MHz (the coarse channel width) — a processing-boundary signature (FFT/polyphase filter bank), not an emitter's. Widening the criterion to any frequency with over 1,000 repeats raises this to 15.83% — but that specific cut is a presentation choice, not a property of the data: hits-per-frequency across the corpus's 250,243 distinct frequencies form a continuous distribution with no bimodal gap (median 9, p99 810). The remaining ~84% (13.7M hits) has no single confirmed cause among three named, undecided hypotheses (grid quantization; a different historical turboSETI version or parameter; a real population of narrowband emitters with drift below the search resolution) — this stays open.
It stops being "rejects signals with no Doppler drift, the classic signature of fixed RFI" (physics, unsupported by the data — the filter does zero work on data generated with a modern turboSETI run in default configuration) and becomes: it rejects a population that is largely specific to this corpus, partially confirmed as an instrument-processing artifact, that modern data generation no longer produces. Correct and sufficient for the current corpus — 56.52% of it is eliminated by this rule, not a marginal effect.
Decision on A as a permanent consensus rule, fixed explicitly: keep A as a permanent rule that essentially never fires (0/25,321 measured) on corpora generated with modern turboSETI defaults — costs nothing where the population doesn't exist, protects against a future corpus with the same characteristic. The conceptually cleaner alternative — A as a corpus-specific rule, gated by a not-yet-existing per-corpus fingerprint architecture — is correct in principle but requires infrastructure that doesn't exist yet; registered as what that future architecture should allow, not implemented now.
Independent of which blocks happened to be mined under the old gate, the corrected cascade was measured over the entire 28,858,709-hit corpus directly (not a biased sample of already-mined candidates):
| Stage | Survivors | % |
|---|---|---|
| Total | 28,858,709 | 100% |
| +A | 12,546,557 | 43.4758% |
| +A+B | 6,328,646 | 21.9298% |
| +A+B+C | 2,453,528 | 8.5019% |
| +A+B+C+D (final) | 113,468 | 0.3932% |
Marginal rejection per filter (share of what reaches each filter that it rejects, not share of the total): A rejects 56.52% of the corpus; B rejects 49.56% of what reaches it (consistent with cadence structure — half of all positions are reference); C rejects 61.23% of what reaches it; D rejects 95.38% of what reaches it — by far the dominant filter on the full corpus.
Sample of 2,000,000 uniform window-start draws (real positions, not a hash simulation): the fraction of 300-hit windows with zero eligible candidates, q, is 0.915237. The real probability of exhausting all 2,000 attempts, q²⁰⁰⁰, is 1.166×10-77 — not a real risk. Expected attempts before finding a candidate: 11.80. (Each window draw is hash(previous_hash, attempt)-derived, not a sequential scan — draws are independent of each other even though the corpus itself is spatially correlated, which affects q but not that independence.) A naive binomial model assuming independent per-position eligibility gives 0.306702 — the real, measured/binomial ratio of 2.98× confirms a prediction made before measuring (spatial correlation of eligibility by scan contiguity).
Economic sizing: 113,468 eligible hits, one per block, at a 300-second target block time = 113,468 × 300s ≈ 394 days per complete pass over the corpus.
Quantifying why a smaller sample (Phase 0, one campaign, one day) measured much higher survival (13.2%) than the full corpus (0.3932%): filter A alone explains a real but partial part of the gap — Phase 0 rejects 0% via A (modern turboSETI) versus 56.52% on the full corpus, a genuine 2.30× effect, leaving most of the 33.6× total gap unexplained by A alone. Filter D's threshold, tested directly (not assumed) by recomputing Phase 0's own 25,321 hits against the full corpus's universe of sources instead of just its own small sample, explains more: survival falls from 13.2% (D against itself) to 3.6452% — a 3.62× reduction, larger than A's effect. Combined (A × D ≈ 8.33×), a residual factor of ~4.03× remains unexplained mechanically — registered as an open residual, not resolved. (This specific cross-corpus D test mixes two campaigns in a way the current architecture doesn't support as a live protocol operation — valid as a one-off diagnostic, not a capability the protocol has today.)
Practical consequence: D gets more selective as the corpus grows — any eligibility estimate made on a small sub-sample of the corpus will systematically overstate real survival, because D has a smaller universe to compare against. This will need to be accounted for whenever future corpora are added.
A methodological error was caught in review before the test ran, and the correction matters more than the result. The first version of this oracle reconstructed the 120 B*-position block list from the pre-reset snapshot cross-referenced against the newly rebuilt production database — new code tested against new code, a self-consistency check, not an independent verification. It would not have caught the failure mode a real oracle exists to catch: a production implementation diverging from the audit's already-validated result, undetected because both sides run the same new code.
Correction: the genuinely independent source was recovered — the audit's own original, unmodified script, reading the chain's pre-reset export and the raw CSV's status column directly, with zero dependency on any of this session's rectification code. Rerun fresh: it reproduced exactly 120/263, and a prediction registered before seeing the result held (corpus-wide B* fraction among hits ≈ corpus-wide B* fraction among mined blocks — consistent with cadence position never having entered the original selection).
Blocking oracle, run against the rebuilt production database:
PASS — this is the genuinely independent oracle the reset's own precaution required, not the self-referential version described above and discarded before it ran.
Verified against the rebuilt production database directly — eight checks, each a direct query against the already-computed result, with no reimplementation of the underlying computation:
NULL — PASS (fail-closed is never ambiguous)._OFF suffix — PASS.PASS — all invariants confirmed. Since the rebuilt database's raw file hash had already been confirmed identical across all four production nodes, this result holds for all four, not only the node it was checked against directly.
Incident 10.1
What happened. A commit expanding the classifier's known-band mask from 9 to 11 bands — classified beforehand as "small change, no production impact" — was pushed to main without the marker this project's CI/CD pipeline requires to skip automatic deployment. The pipeline deploys to all four production nodes automatically after tests pass unless that marker is present — a real, documented mechanism, but not something any task list in this audit had flagged as relevant to a change of this kind. Result: the two new bands went live on all four nodes within minutes of the commit, before the classifier retrain that was supposed to accompany it, and before the planned reset — the reverse of the originally specified order.
The attempt to intervene, and why it didn't work. On noticing the problem, with tests already green and deployment about to run, an attempt was made to cancel the running CI job via the CI platform's API — the endpoint returned 404; it doesn't exist on this instance. A second attempt to merely check the job's status was blocked by this session's own automated-mode permission guard. No attempt was made to work around that block — it was reported to the operator instead.
Decision taken (operator): let it run, retrain immediately afterward, don't revert. Reverting would have been a second live production change to undo the first, with CI already in flight and no clean cancellation path — more risk, not less. The damage was temporary by construction: it affected the classifier's declared confidence score on new blocks until the retrain landed, and the planned reset would invalidate those blocks regardless. On a testnet with no monetary value, the real cost of the temporary mismatch was low.
A commit to this repository's main branch is a production action, not a step that precedes one. This wasn't explicit in any task list before this happened — which is exactly why an item classified as "small, no production involved" ended up changing four live nodes on its own, ahead of the retrain it was supposed to ship with. Going forward, any code change that shouldn't go immediately to production needs the skip-deployment marker (or a branch, not main) by default — not as a case-by-case exception to remember.
The retrain itself completed the same day: ROC-AUC 0.9623; the known-band feature's importance in the retrained model remained marginal (0.0045) even with the corrected mask — consistent with two separate, independently-supported conclusions, not one implying the other: (a) this is not a mask failure — with roughly 99% of hits falling inside some known band, the feature is nearly constant in training, so low importance is expected regardless of whether the mask is accurate; (b) the classifier itself plays no role in consensus at all — the four hard-rejection rules do the real filtering, and no node ever recomputes the declared confidence score to validate a received block. One node in the production fleet ran with no classifier model file present at all during this incident — confirmed, by direct evidence rather than inference, that this was inert rather than a gap: that node never mines (so never needs to compute a fresh score) and never recomputes a received block's score to validate it (so the model's absence was never exercised in production).
Incident 10.2
What happened. The first run of the full-corpus eligibility measurement (28,858,709 hits) used the unmodified production import code to build a separate index. The cadence-coincidence filter (C) stage became visibly slow; a live process inspection confirmed it was genuinely working, not stuck — roughly 40 minutes of CPU into that stage, processing a single cadence group of 22,215 hits, at an observed rate of ~601 rows/second. Extrapolated: ~13 hours for the full corpus.
Why this mattered beyond being slow: this is the exact code the actual production reset would have to run. Without this fix, the reset itself would have carried the same ~13-hour cost — discovered only after starting, not before.
Cause: filter C's comparison loop checked each row against every reference-position hit in its cadence group, with no pruning — O(group size) per row, O(group size²) per group. No test before this had ever exercised this at scale: unit tests use small synthetic groups, and the only prior real-data run used a single cadence with 25,321 hits total. The full corpus has individual cadence groups with tens of thousands of hits each.
Fix: the original specification for this exact optimization already existed in this project's own task list — "sort by frequency, sliding window — avoids quadratic blowup" — but had never actually been implemented in C's real loop. Implemented now: within each cadence group, reference-position hits are sorted by frequency once; each row's search radius is bounded by that specific group's own worst case (its own maximum drift times its own maximum time span — a real, derived upper bound, never a guessed global constant), then each row does a narrow binary-search lookup instead of a full group scan.
Equivalence verified before trusting the fix, not just argued: 39,000 synthetic hits across 3 deliberately varied groups, old algorithm (extracted unmodified from the prior commit) and new algorithm run against the same data, every output column compared row by row by a stable key. Zero discrepancies across 39,000 rows. A complexity-regression test was added: two synthetic group sizes (9,000 and 18,000 hits), asserting that doubling the size does not multiply the time by anywhere near 4× (what a reintroduced quadratic algorithm would produce) — a ratio test rather than an absolute time limit, so it isn't sensitive to which machine runs it.
This was the third time in the same investigation that a specification already written down was never actually verified against the implementation — the scan-position deduplication bug (section 4), the cadence-boundary detection bug (also section 4), and this one. In every case, the correct decision had already been written in the right document before the bug appeared. The lesson is not "write better specifications" — it's that a specification was never checked against the code after being written; it stayed a documented intention, not a confirmed property of the implementation. This is the direct argument for treating any algorithm with a specified shape (ordering, windowing, a bounded search) as needing an automated growth-shape test, not just a read-through, before being considered done.
The classifier (RandomForestClassifier) scores each hit on four features: frequency, drift rate, SNR, and whether the frequency falls inside a band already known to be crowded with terrestrial transmitters. Cadence structure and cross-target frequency coincidence — the two things that actually discriminate a real candidate — are never features of the classifier; they are computed separately, only ever used to build training labels, and never reach the live scoring decision.
The known-band mask was originally three narrow bands, two of which don't even intersect this corpus's frequency range (1088–1909 MHz) at all — making the classifier's one piece of contextual information practically dead. The correction cannot be "center bands on the already-known survivors" — that would be a circular exclusion list, built to eliminate known cases, indefensible against "why exactly here?" and useless on new data. The corrected mask instead uses published spectrum allocations covering the corpus's full range: Mode S/ADS-B, GNSS, ARSR radar, Inmarsat, RNSS/GPS/GLONASS/Galileo, Iridium, MetSat, PCS.
Verified against primary sources directly, with the same discipline as the rest of this audit: the original nine bands had been cited from memory; independently confirmed against NTIA Spectrum Compendium documents, the FCC Table of Frequency Allocations, and a published paper on Iridium interference in radio astronomy. The verification caught a real citation error, not a cosmetic adjustment: MetSat is 1675–1710 MHz per NTIA, not 1670–1710 as originally cited — a 5 MHz range that concretely reclassifies one specific block, verified directly rather than assumed.
Text-based parsing of the frequency field versus the float-multiply approach it replaces: zero divergences across all 28,858,709 rows. This is a definitional fix, not a correction of already-corrupted data — the existing approach happens to be safe for this specific corpus; the change is about robustness against future data with different formatting, not a problem that has already manifested.
100% of the corpus has uniform 12-decimal-place precision in the MJD field (zero cases of scientific notation, zero malformed values) — unlike the frequency field, which has variable width. The minimum non-zero gap between scans of the same obs_run is 322 seconds; the maximum span within a single obs_run is 21,321 seconds (5.9 hours) — revealing that obs_run is not itself a single cadence unit; it can contain several independent cadences.
A hit's authenticity is checkable only against the published CSV, never against the original telescope observation. The raw observation files this CSV was exported from aren't part of this corpus at all; no node confirms the CSV matches what the telescope actually recorded, only that a given row exists in the canonical CSV. This is mitigated, not eliminated, by the CSV's own public provenance — anyone can independently re-derive it from the same published archive.