Operations first
Is the pipeline working?
Pipeline execution, accepted prediction state, and browser delivery are reported separately.
Latest pipeline attempt
LoadingSource freshness
Why matches are blocked
One primary group per blocked row; open a match for its exact audit reasons.
Capital and exposure
Data integrity boundary
“Feature complete” means required values are present. It does not certify that upstream match history, duplicate quarantine, or train/serve parity passed an independent audit. Exact lineage remains visible on every match.
Accepted prediction run
Current slate by tournament
Only rows from the last successful prediction-bearing run are considered.
Accepted-slate eligibility funnel
One cohortLoading the accepted run…
Decision inputs valid
Pre-start rows with complete metadata, two-sided prices, model output, and an exact feature snapshot ID.
Blocked / incomplete0
These are monitored but not actionable. The reason is stated on each row.
Started / expired0
Retained for audit only. They are never counted as live or bettable.
Diagnostic, not promotion evidence
Model performance laboratory
production.evaluation.ledger during durable publication and withheld if their sync generation disagrees with the manifest. Tour filters are also server-materialized: ATP includes ATP-level, Masters, and Grand Slam rows; Challenger and ITF remain separate. The linked Markdown report is a dated full ledger snapshot (live metrics plus offline experiments) and may lag this manifest-pinned table. Green and red compare like-for-like rows against the market when the cohort is shared; ROI uses zero as break-even. Per-model cohorts retain their own n and are exploratory, not a fair ranking. Kalshi ROI uses raw buy asks without de-vigging, but counts a row only after exact entry rules, entry-time fee evidence, and terminal exchange settlement values are all hash-verified. It is forward-only from the displayed rule-capture date.
Parallel populations—not one sample
What each performance number counts
Prediction cohorts measure model quality. Counterfactual bets replay a rule on those predictions. Placed-bet outcomes measure the paper ledger; attribution-uncertain recoveries release capital and preserve P&L but never become model evidence.
Open dated full ledger snapshot This report may lag the accepted sync shown above.
Scalar summaries
Quality scores at a glance
These are scorecards, not diagnostic curves. Use ROC and calibration below to inspect model behavior.
Version-aware comparison · cohort labels explicit
Versioned ROC comparison
Choose any model lines below. Live-forward curves use the growing exact-settled sample and show their own n. With sealed replay enabled, every replay v2, market, and production-reference curve uses the same fixed 629-match cohort. Dots are the exact authoritative thresholds; smoothing only rounds the visual guide and never changes AUC or the underlying points.
One curve · every exact threshold
ROC threshold detail
Inspect any available production, v2 live-forward, or sealed-replay curve and its full threshold table. Use the versioned comparison above to overlay selected lines. The diagonal is random ranking; AUC summarizes the area under the curve.
Reliability diagram
Calibration
Dots above the diagonal are underconfident; below are overconfident. Empty bins are omitted.
Scalar comparison—not a diagnostic curve
Metric explorer
Candidate lane · never actionable
Ratings v2 evidence
Candidate reliability
Calibration
How to read shadows and ROI
Each shadow variant uses one deterministic opening observation per match and model version, joined to the operational opening feature snapshot. Hourly repeats do not increase n. “Actual paper” performance remains in Paper bets; both flat figures here are counterfactual cohort replays. Bovada uses the offered decimal odds and raw break-even hurdle. Kalshi uses each side’s raw ask as the hurdle with no de-vigging, then requires the exact per-market rule envelope, entry-time fee evidence, and exchange terminal value before scoring P&L. Generic match winners and generic filings are not settlement authority. Treat n below 250 as exploratory even when ROI is green.
Paper ledger
Bets and realized P&L
Win, loss, void, cancelled, and pending states are handled explicitly.
| Placed | Match / side | Price | Stake | Edge | Status | P&L |
|---|
Settled history
Prediction results
| Match date | Event | Match | Winner | Score | NN | XGB | Market | Market path |
|---|
Manifest-pinned registry view
Models & versions
What is promoted, what is only being measured, and what changed across releases.
| Model | Version | Lifecycle | Features | Probability | Released | Artifact | Notes |
|---|
Audit surface
Runs and mirror state
Status, errors, and stage counts remain visible even when a run produced no predictions.
| Run | Started | Status | Odds | Features | Predictions | Bets | Settled | Capital / exposure | Post-run stages | Error |
|---|