● An original study · Ten AI models · Seven classic money tests
Tails, you lose.
Ten AI models took the experiments that made behavioral economics famous. They beat us on sunk costs, lost to a random number, and took a losing bet the moment it was called a trade.
Fig. 00 · Petty cash · 0 coins
Heads 0 · Tails 0
Shove them around · Click a coin to flip it
Fair coin · p = 0.5
Scroll to open the ledger
The books, briefly.
Every figure on this page is computed live from the raw trial data. Nothing is hand-typed.
Entry A · Models audited
10
Open-weight models from Alibaba, Meta, Mistral, Google and Microsoft, 0.6 to 14 billion parameters.
Entry B · Decisions logged
31,500
Each one a fresh conversation. Every model saw one version of one problem at a time, like a human subject.
Entry C · Biases tested
7
Framing, loss aversion, anchoring, sunk cost, the disposition effect, mental accounting, the certainty effect.
Entry D · API spend
$0.00
Every trial ran locally, overnight, on one consumer graphics card.
Bottom line
Loading the ledger…
Take the test they took.
Seven sealed entries. You get one randomly assigned version of each problem, exactly like the models did. The name of each bias stays sealed until you answer.
The audit.
Bias quotient: each model's bias divided by the human bias on the same test. 0 is perfectly consistent. 1 is exactly as biased as people. Hover any cell.
Five findings.
Each finding compares the same models under two conditions. Lines show the change. Hover any row for exact figures.
Reconciliation
Answer the seven entries to see who you decide like.
Read the full paper.
Method, statistics, limitations and every prompt used. The code and raw data are public.
Working paper · PDF
Tails, You Lose: Do Open-Weight Language Models Inherit Human Biases in Financial Decisions?
Daniel Yerushalmi · 2026
Download ↓Design
Between-subjects. Every trial is a fresh conversation with one version of one problem. Option order is shuffled every trial so a model's taste for "Option A" cannot pose as a preference.
Wording
Each test exists twice: the famous textbook wording, and a new financial scenario with identical structure. Comparing the two shows whether behavior on famous problems carries over to new ones. Mostly, it did not.
Modes
Fast (answer only) and thinks-first (writes its reasoning before answering), plus a financial-advisor persona. 30 trials per cell, temperature 0.8.
Built with
Extensive AI assistance (Anthropic's Claude) under the author's direction. Every prompt, raw model answer and line of analysis is public, so every number can be checked.