Promises, checked
Peer reviewed · two independent checks re-derived every figure from the raw inputs and reached the same conclusions, 20 September 2026

This page in one lineThis page checks whether stated chances come true as often as they should, for the betting crowd in prediction markets and for Charlie’s own confidence scores, so you know how much to trust a number like “70%”.

Does 70% Mean 70%?

Across 324 large prediction markets that have finished, when the crowd said a thing had a 97% chance a week before the end it happened 94% of the time. Overall the crowd’s week-before prices were 67% better than just saying 50% every time, but 70% of these markets were near-certain either way; on the genuinely uncertain ones the crowd was only 18% better than a coin flip. Charlie’s own calls won 65% of the time, and the highest-confidence quarter won 77% against 70% for the lowest, but the middle groups were out of order, so the score is only a rough guide. Charlie’s separate forecasting tool for Bitcoin price levels has 231 forecasts still waiting to finish; the first can be graded from 2026-12-31.

Finished markets checked324The largest finished markets of the last year; the smallest had $13 million of bets
Crowd’s score on uncertain markets18% betterthan a coin flip, on the 96 markets priced between 10% and 90%
Charlie’s highest-confidence quarter77% wonversus 70% for the lowest quarter
Bitcoin forecasts still waiting for a grade2310 finished so far; first grade 2026-12-31

When the crowd said 70%, did it happen 70% of the time?

Every dot is a group of finished prediction markets. Left to right is the chance the crowd’s price stated one week before the market closed. Up and down is how often the thing actually happened. If promises were perfect, every dot would sit on the dotted line. Dots above the line mean the crowd was too gloomy; dots below mean too confident. Bigger dots hold more markets.

0%0%20%20%40%40%60%60%80%80%100%100%The promise line: said 70%, happened 70%21015961618197618What was promisedWhat actually happened

324 markets. The label on each dot is how many markets it holds.

Charlie’s readThe crowd is close to honest. Among groups with at least 15 markets, its biggest miss was the 40–50% group, which said 46% and saw 31% happen, though that is only 16 markets. Things it called near-certain (97%) happened 94% of the time. Long shots (outcomes the crowd thought very unlikely) priced at 1% came in 3% of the time: rare, as promised, but about 3.5 times more often than the price said.

What this does not mean. A dot near the line does not mean the crowd predicted each event. It means that across many events, its stated chances added up honestly.

How much of the crowd’s good score comes from easy calls?

Most finished markets end up near 0% or near 100% a week before they close, because by then everyone knows the answer. Getting those right is easy. So the crowd’s score is shown three ways: against a coin flip on every market, against someone who simply knew how often “yes” wins in general (23% of the time here), and on the genuinely uncertain markets only, those priced between 10% and 90%.

+0%+20%+40%+60%67% betterAgainst a coinflip, all markets53% betterAgainst knowingthe usual yes-rate18% betterUncertain markets only,against a coin flip

70% of the 324 markets were priced under 10% or over 90% a week out. The uncertain set holds 96 markets.

Charlie’s readThe headline score of 67% is real but flattering. On the hard markets the crowd was 18% better than a coin flip: still useful, and still honest on average, but not magic.

What this does not mean. A lower score on hard markets is not a failure. Hard markets are hard for everyone; the point is to know how big the crowd’s advantage really is.

Does the crowd get better as the deadline approaches?

Mostly the same markets (286 of the 324 had a price a month out), but now the stated chance is taken a month before the close instead of a week. Markets learn as news arrives, so the month-before dots should be a little further from the line.

0%0%20%20%40%40%60%60%80%80%100%100%The promise line: said 70%, happened 70%185151110131816936What was promisedWhat actually happened

286 markets with a price available a month before the end.

Charlie’s readA month out, the crowd’s prices were 66% better than a coin flip; a week out, 67% better. The two are almost the same, which means the last three weeks of news moved these prices very little.

What this does not mean. Better later is normal and not a flaw. The point is to know how much a price should be trusted at each distance from the deadline.

Did Charlie’s higher confidence mean more wins?

The Alpha Tracker is Charlie’s public list of short-term trading calls, each saying a small coin will rise or fall over the next hours or days. Every call gets a confidence score before it is published. The 255 finished calls are split into four equal groups by that score. If the score means anything, the bars should climb from left to right.

+0%+20%+40%+60%70% of 64Lowestconfidence53% of 64Low60% of 63High77% of 64Highestconfidence

All calls together: 65% won. The calls were published on only 16 separate days, so every group contains calls made in the same market conditions.

Charlie’s readThe bars do not climb cleanly. Higher confidence did not reliably mean more wins. The score is not a promise of a percentage, so it is judged on order, not on matching a number.

What this does not mean. With calls bunched on a few days, a group can look strong because it happened to be published on a good day.

Did higher confidence mean bigger wins?

Winning more often is one thing. This chart shows the typical size of the result in each confidence group, in percent.

+0%+1%+2%+1.7%Lowestconfidence+0.5%Low+1.8%High+2.6%Highestconfidence

Rank agreement, a number from -1 to +1 saying whether higher scores went with better results, where 0 means no link: +0.09.

Charlie’s readThe rank agreement of +0.09 is close to zero: the score did not sort the results by size.

What this does not mean. Typical result here is the middle call, not the average, so one giant winner cannot carry a group.

Which part of Charlie’s score did the work?

The confidence score is built from three parts: the shape of the price chart, whether there was a news trigger, and where money was flowing. For each part, calls that scored higher on it are compared with calls that scored lower; the number in brackets is how many calls fell in each group, and when most calls share one value the groups are far from equal.

Chart shape · upper half (128)62% wonNews trigger · at or above the usual score (232)64% wonMoney flow · upper half (128)69% wonChart shape · lower half (127)68% wonNews trigger · below it (23)74% wonMoney flow · lower half (127)61% won+0%+20%+40%+60%

A part that matters shows a clear gap between its two groups, with the higher-scoring group winning more. The number in brackets is how many calls are in the group.

Charlie’s read“News trigger” made the biggest difference, -10 points, but in the wrong direction: calls that scored higher on it won less often. Not every part pointed the right way, so the overall score is carrying parts that hurt it in this period. Groups with only a few dozen calls are easy to fool.

What this does not mean. Which part works can change with the market. This is a report card for one period, not a permanent ranking.

What is still waiting to be graded?

Charlie also runs a separate forecasting tool that states a chance of Bitcoin reaching certain price levels by certain dates, beside the chance the betting price works out to. None of those forecasts has finished yet, so they cannot be graded here. The chart shows when they will be. All of them finish on the same day, so the first grade will be one batch judged on one day, not many separate tests.

050100150200231Dec 2026

231 forecasts open. This page will grade them as they finish and show the result the same way as the charts above.

Charlie’s readSaying this plainly matters. A forecast record that only shows finished wins is not a record. The first grade arrives 2026-12-31.

What this does not mean. An ungraded forecast is not a failed one. It is just not evidence yet.

What should you do with this?

Two kinds of reader use this page. One runs money for other people and has rules to follow. The other is deciding about their own savings. The same evidence leads to different actions.

If you run money for others

Funds, trading teams, the people who manage a company’s cash, research teams.

  • You rely on prediction-market prices in your own workA week-before price is a usable probability on average. Trust it most on markets already near 0% or 100%, and treat mid-range prices as a rough guide, not a precise one.
  • You are shown a confidence score by any providerAsk for this chart: win rate by confidence quarter. If the bars do not climb, the score is decoration.
  • You want to grade a forecasterInsist on the full record including unfinished forecasts, and on the number of separate days behind it.

If it is your own money

Anyone deciding what to do with their own savings.

  • A market says 90% and you want to bet on the likely outcomeNear-certain things did come in about as often as promised here, but you win very little when right and lose a lot on the rare miss.
  • A market says a few percent and you like the unlikely outcomeVery unlikely outcomes here came in a little more often than the price said, but they still lost almost every time. They are not free money.
  • You follow Charlie’s callsTreat the confidence score as a rough guide only: the groups did not climb in order. The record is only a few weeks old.

The conclusion

What this page is

A check of whether stated chances came true as often as promised, for the prediction-market crowd and for Charlie’s own confidence scores.

Why it matters

A percentage is only useful if 70% actually means 70%. Most forecasters are never checked this way, so their numbers sound precise and mean little.

How it is useful

It tells you how far to trust a market price on easy and on hard questions, and whether Charlie’s confidence score is worth reading.

The crowd kept its promises well on average, but most of its 67% score came from easy calls; on uncertain markets it was 18% better than a coin flip. Charlie’s confidence score did not sort calls cleanly. Charlie’s Bitcoin price-level forecasting tool has 231 forecasts still waiting for their grade.

Words used on this page

Every technical word above is explained again here, in plain English.

  • Prediction marketA place where people bet real money on whether something will happen. The price is the crowd’s stated chance.
  • Stated chanceThe probability a price or a score is claiming. A price of 70 cents on a yes-or-no bet means a 70% stated chance.
  • Promise lineThe dotted diagonal. A dot on it means the thing happened exactly as often as the stated chance said.
  • Better than a coin flipHow much lower the forecaster’s error was than someone who said 50% every time. 0% is no better; 100% would be perfect.
  • Confidence scoreA number Charlie attaches to each call before publishing it, built from the chart, the news and the money flow.
  • Alpha TrackerCharlie’s public list of short-term trading calls on small coins, each logged before its outcome is known.
  • CallOne published prediction that a coin will rise or fall over the next hours or days.
  • Long shotAn outcome the crowd thinks is very unlikely, priced at a few percent.
  • Uncertain marketsMarkets priced between 10% and 90% a week before the close, where the answer was still genuinely in doubt.
  • Forecasting toolCharlie’s separate model that states a chance of Bitcoin reaching a given price level by a given date.
  • Money flowWhere money was moving into or out of a coin around the time of a call, one of the three parts of the confidence score.
  • Quarter or groupThe calls split into four equal piles by confidence score.
  • Rank agreementA number from -1 to +1 saying whether higher scores went with better results. 0 means no link.
  • GradedA forecast whose deadline has passed, so it can be marked right or wrong.

Where this page could be wrong

  • The prediction markets here are the largest finished ones from the last year; the smallest carried $13 million of bets. Small markets behave worse and are not included.
  • The middle groups on the dot charts hold only a handful of markets each, so one market can move a dot by several points.
  • A market’s stated chance is read from its price one week and one month before the end. Prices in between are not used.
  • Charlie’s calls were published on a small number of days, so the four confidence groups share market conditions.
  • The confidence score was never meant to be a percentage. It is judged on order, not on matching a number.
  • All of this is about what has already finished. It says how honest past promises were, not whether the next one will be.