Examples

How an AI can use calibrated probabilities and evidence strength.

One of these has now resolved — shown first. A single resolved call isn’t a track record. It shows how an agent can combine a calibrated probability with group-conditional evidence strength under an explicit decision policy.

Resolved · 16 May 2026

Where should I host my Eurovision party?

A Denmark-themed party: commit to a premium venue (worth +€500 if Denmark finishes top-5, −€500 if not), or host at home. Before the final — forecast pulled from the production API on 2026-05-12, four days out:

Claude + Internet → book the venue.

Polymarket’s 66.5% is above the 50% breakeven, so expected value looks positive (+€165).

Claude + Credence → host at home.

Credence’s calibrated probability was actually higher — 72.7% — so the point estimate still said book. Its evidence strength was 4.80, fitted from comparable historical forecasts rather than this market. Under an illustrative conservative policy that required stronger group-level evidence before an irreversible commitment, the call was to pass.

Why it mattered: Credence wasn’t “more right” about the odds — in fact it leaned more toward Denmark than the market did. The evidence-strength value described the historical cohort assigned to the forecast, not latent precision for Denmark. Claude used that group-level fact in a conservative commitment rule.

What happened — 16 May 2026: Denmark finished 7th, outside the top five (Bulgaria, Israel, Romania, Australia, Italy). The venue bet would have lost €500. Claude + Credence’s call: €0.

Driving market: Denmark top-5 at Eurovision 2026. Market price 66.5% · calibrated probability 72.7% · group-conditional evidence strength 4.80, pulled from the production API 2026-05-12 — before the 16 May final. Outcome real (Denmark 7th); the venue payoff and decision policy are illustrative. Not advice.

Show Claude’s math

EV (venue − home | p) = 1000 p − 500 ; breakeven p = 50%.

Market mean 0.665 → +€165 (venue). Credence mean 0.727 → +€227 (still venue).

Evidence strength 4.80 describes the assigned historical cohort; it does not create a market-specific precision estimate.

Illustrative policy: do not make the irreversible venue commitment when the assigned cohort’s evidence strength does not clear the decision-maker’s pre-set floor → home.

Assumption: this group-level risk policy. Credence supplies the probability object; it does not prescribe the action or the floor.

Realized 2026-05-16: Denmark 7th, not top-5 → venue −€500, home €0.

Should I rent a tent for this weekend’s event?

An outdoor celebration spans two days. One wet day is fine; rain on both is a washout that forfeits $40,000 in deposits. A tent costs $6,000 and removes the risk.

Claude + Internet → skip the tent.

Treating the 35% daily-rain price as fixed, the washout chance is 12.3% → expected loss $4,900, cheaper than the tent.

Claude + Credence → rent the tent.

Same 35% calibrated probability, plus group-conditional evidence strength of 4.0. Under an illustrative downside policy that required stronger comparable-cohort evidence before risking the deposits, the agent bought protection.

Why it flipped: The calibrated probability is identical to the market’s. The agent’s own policy uses group-conditional evidence strength as a class-level risk input; it does not treat it as a market-specific uncertainty estimate.

Driving market: rain at the venue on an event-weekend day. Market price 35.0% · calibrated probability 35.0% · group-conditional evidence strength 4.0. Illustrative construction; not advice.

Show Claude’s math

Washout = rain both days = p². Tent $6,000; washout loss $40,000; breakeven E[p²] = 0.15.

Plug-in: 0.35² = 0.1225 → EV (skip) = −$4,900 → SKIP.

Group-conditional evidence strength = 4.0. In this illustration, the buyer’s pre-set downside policy requires stronger comparable-cohort evidence before accepting the $40,000 loss exposure → RENT.

Assumption: this evidence-strength floor is the buyer’s policy. It is not an API-prescribed threshold or a market-specific precision claim.

Should I book refundable or non-refundable travel?

A big trip abroad. Non-refundable bookings are cheaper but forfeit $20,000 if the trip is derailed; fully flexible booking costs $6,000 more. The trip is derailed if any of four unrelated risks hits — a strike, a storm, an entry-rule change, or a transport disruption.

Claude + Internet → book non-refundable.

The four market prices compose to a 27.7% chance of derailment — below the 30% where flexibility pays for itself.

Claude + Credence → book refundable.

Each risk is calibrated a little higher; none decisive alone, but compounded across four they lift the derailment chance to 33.8%, past breakeven.

Why it flipped: No single market moves the decision. Small calibration corrections compound across unrelated markets into one that does.

Driving markets: four independent risks (strike / storm / entry policy / transport). Joint derailment: 27.7% market vs 33.8% calibrated. Illustrative construction; not advice.

Show Claude’s math

Derailed = any of four = 1 − ∏(1 − pᵢ); breakeven 30%.

Market: 1 − (.92)(.88)(.95)(.94) = 27.7% → EV (rigid) −$5,541 vs flex −$6,000 → RIGID.

Credence: 1 − (.90)(.86)(.93)(.92) = 33.8% → EV (rigid) −$6,755 → FLEX.

No single risk alone crosses 30% (each lands near 29%); only the four compounded do.

Assumption: the four risks are modeled as independent.

Should I realize capital gains this year or next year?

A near-retiree holds a concentrated position with a large unrealized gain. Defer 18 months hoping a tax change repeals a surtax, or realize now and diversify.

Claude + Internet → defer.

At the market’s 46.5% for Senate control, expected value slightly favors waiting (+$277).

Claude + Credence → realize now.

Calibrated to 53.8%, the repeal path is less likely than the market implies; expected value flips to favor realizing.

How to read it: repeal of the surtax runs through a Republican Senate, so a higher chance of Democratic control means a lower chance of repeal — Credence’s higher 53.8% therefore pushes toward realizing now, not deferring.

Why it flipped: Here it’s the calibrated mean that moves the decision — across a real economic threshold the market was sitting on.

Driving market: Democratic Senate control after the 2026 midterms. Market price 46.5% · calibrated probability 53.8%. Inputs pulled from the production API, 2026-05-12. Illustrative; not advice.

Show Claude’s math

P(repeal) = 0.45·(1 − p_dem) + 0.02·p_dem.

EV (defer − realize) = $15,200 · P(repeal) + $5,769 − $9,293.

Market p_dem 0.465 → P(repeal) 25.005% → EV = $15,200(.25005) + $5,769 − $9,293 ≈ +$277 → DEFER.

Credence p_dem 0.538 → P(repeal) 21.866% → EV = $15,200(.21866) + $5,769 − $9,293 ≈ −$200 → REALIZE.

Breakeven p_dem ≈ 50.7%; the calibrated probability is above it.

Assumption: repeal is worth $15,200 if it occurs.

How the agent decides.

Credence supplies a calibrated point forecast plus group-conditional evidence strength. For expected-value choices, an agent can use the point forecast. For downside-sensitive choices, it can apply a predeclared risk policy that treats the evidence strength of comparable historical cohorts as a class-level input.

There are two ways the current probability object changes the calls above:

  • the corrected mean crosses a decision threshold — on one market, or by compounding across several
  • a downside policy uses group-conditional evidence strength without pretending to know one market’s latent precision

Credence does not prescribe either the decision rule or an evidence-strength floor. Your AI does the reasoning; Credence supplies the traceable probability object behind it.

Illustrative demonstrations of how a probability object API changes a downstream decision — not tax, legal, investment, or event-planning advice. Figures marked ‘production API’ were pulled 2026-05-12; others are illustrative constructions. The Eurovision outcome (Denmark 7th, 16 May 2026) is a real resolved result and the forecast was a real production-API pull; the venue payoff is an illustrative scenario.