Methodology
We calibrate prediction-market probabilities against outcomes.
A market price folds trades and available information into one number. It does not say how comparable forecasts held up historically, or whether recurring patterns biased the price.
Credence returns a calibrated probability and group-conditional evidence strength. Both are validated across groups of forecasts, not inferred as latent precision for a single market.
Calibrated means real-world.
A probability is calibrated across many forecasts: events assigned 70% should resolve YES about 70% of the time. No single terminal outcome can establish calibration by itself.
We don’t aim to beat the market on every question; we aim for Credence’s 70% forecasts to mean 70% across comparable cases.
Two claims, with the same validation ceiling.
The current response separates the point forecast from the historical evidence around forecasts like it:
- Calibrated probability — our point forecast for the real-world outcome.
- Evidence strength — a summary fitted from held-out cohorts of comparable forecasts.
Evidence strength can vary with the historical cohort assigned to a forecast. It is not an effective sample count for the returned market, and it does not identify that market’s latent precision.
The honest reading
Markets like this one.
Not: this market contains a measurable amount of latent evidence.
Evidence strength has a hard identifiability limit.
A market supplies one terminal label. From that label, per-market latent dispersion is provably unidentifiable: the likelihood contains information about the mean, but none about a separate dispersion parameter for that market.
The strongest honest label-validated claim is therefore group-conditional. Evidence strength describes how held-out cohorts of comparable forecasts behaved; it never promises per-market latent precision.
Tested against reality, out of sample.
Numbers are checked against outcomes the model did not train on for that test fold. We score with proper scoring rules, measure calibration directly, and evaluate evidence-strength behavior at the cohort level.
No look-ahead. No grading our own homework.
Reproducible, versioned, always improving.
Every response is pinned to a snapshot ID, model version, release, methodology version, and source time so the answer can be traced to the versioned inputs and runtime that produced it.
The research moves fast, but changes ship behind a stable contract, are gated before release, and are tagged on every response. You can tell which methodology answered without relying on unstated semantics.