An original study · Ten AI models · Seven classic money tests

Tails, you lose.

Ten AI models took the experiments that made behavioral economics famous. They beat us on sunk costs, lost to a random number, and took a losing bet the moment it was called a trade.

++++

Fig. 00 · Petty cash · 0 coins

Heads 0 · Tails 0

Shove them around · Click a coin to flip it

Fair coin · p = 0.5

Scroll to open the ledger

The books, briefly.

Every figure on this page is computed live from the raw trial data. Nothing is hand-typed.

Entry A · Models audited

10

Open-weight models from Alibaba, Meta, Mistral, Google and Microsoft, 0.6 to 14 billion parameters.

Entry B · Decisions logged

31,500

Each one a fresh conversation. Every model saw one version of one problem at a time, like a human subject.

Entry C · Biases tested

7

Framing, loss aversion, anchoring, sunk cost, the disposition effect, mental accounting, the certainty effect.

Entry D · API spend

$0.00

Every trial ran locally, overnight, on one consumer graphics card.

Bottom line

Loading the ledger…

Take the test they took.

Seven sealed entries. You get one randomly assigned version of each problem, exactly like the models did. The name of each bias stays sealed until you answer.

The audit.

Bias quotient: each model's bias divided by the human bias on the same test. 0 is perfectly consistent. 1 is exactly as biased as people. Hover any cell.

Answering
Wording
Persona

Five findings.

Each finding compares the same models under two conditions. Lines show the change. Hover any row for exact figures.

Reconciliation

Answer the seven entries to see who you decide like.

Read the full paper.

Method, statistics, limitations and every prompt used. The code and raw data are public.

Working paper · PDF

Tails, You Lose: Do Open-Weight Language Models Inherit Human Biases in Financial Decisions?

Daniel Yerushalmi · 2026

Download

Design

Between-subjects. Every trial is a fresh conversation with one version of one problem. Option order is shuffled every trial so a model's taste for "Option A" cannot pose as a preference.

Wording

Each test exists twice: the famous textbook wording, and a new financial scenario with identical structure. Comparing the two shows whether behavior on famous problems carries over to new ones. Mostly, it did not.

Modes

Fast (answer only) and thinks-first (writes its reasoning before answering), plus a financial-advisor persona. 30 trials per cell, temperature 0.8.

Built with

Extensive AI assistance (Anthropic's Claude) under the author's direction. Every prompt, raw model answer and line of analysis is public, so every number can be checked.