Pump. Bank.
Or pop.
A five-minute behavioural measure of risk appetite: the same laboratory task used in over two decades of published psychology research, rebuilt for the browser. Free to run, free to share.
One balloon. One decision, over and over.
Every pump is a gamble: the balloon is worth more, and it is closer to bursting. Nobody tells you where the limit is; that is the point.
Pump
Each pump inflates the balloon and adds money to a temporary pot for that balloon. Bigger balloon, bigger pot, and a bigger chance of losing it.
Bank
Stop at any moment and collect the pot into your permanent bank. Banked money is safe forever, but you'll never know how much you left on the table.
Pop
Every balloon has a hidden burst point: some pop on the first pump, some survive a hundred. Pop it, and that balloon's pot vanishes.
A real scorecard, not a quiz result.
Scores use the standard research metric, adjusted average pumps, with a balloon-by-balloon chart, an approximate percentile against published adult samples, and a CSV export of every trial.
Reading a BART score
What the numbers mean, where the cut-offs come from, and (just as important) what a balloon task can't tell you.
The headline number: adjusted average pumps
The score that BART research is built on is the adjusted average number of pumps: the mean number of pumps on balloons that were banked, ignoring balloons that popped. Popped balloons are excluded because a burst censors the measurement: we never learn how far the person would have pushed that balloon, only where chance stopped them (Lejuez et al., 2002).
More pumps means more risk accepted per decision. Under the standard settings (where each balloon bursts somewhere between the 1st and 128th pump, uniformly at random), the mathematically optimal strategy is to aim for about 64 pumps per balloon. Almost everyone stops far short of that, which is exactly what makes the task informative: it measures the gap between the risk people could profitably take and the risk they are willing to take.
Typical ranges
Across published healthy-adult samples on the standard 30-balloon task, adjusted averages mostly land between the mid-20s and mid-40s. As a rough map:
| Adjusted average | Band | Typical reading |
|---|---|---|
| under ~25 | Cautious | Banks early and often. Consistent, low-variance earnings; leaves substantial value unclaimed. Common in risk-averse and loss-attentive profiles. |
| ~25 – 45 | Typical | The bulk of adult samples. Balances the pull of a growing pot against the sting of occasional pops. |
| over ~45 | Risk-seeking | Pushes balloons hard and absorbs more pops. Higher scores correlate with sensation-seeking and real-world risk behaviours in the research literature. |
Bands are approximate, assembled from published means and standard deviations for healthy adults on the 30-balloon, 128-pump version. They are context, not clinical cut-offs, and they only apply when the task is run with standard settings.
The supporting numbers
- Balloons popped. Raw appetite for pushing past comfort. Popping 8–12 of 30 is unremarkable; popping none suggests strong loss aversion.
- Total earned. The outcome of the strategy, luck included. Two people with the same adjusted average can earn quite different amounts.
- Early vs. late pumping. Comparing the first and final thirds of the run shows adaptation. Most people drift slightly upward as they calibrate; a sharp drop after a pop, followed by recovery, is a healthy learning signature.
- Pump rhythm. Median time between pumps. A steady rhythm suggests a pre-committed target; long hesitations near the end of a balloon suggest decision conflict at the margin.
What a BART score is not
A single five-minute task is a snapshot of behaviour under one specific kind of uncertainty: escalating gains with sudden total loss. It is not a personality profile, an integrity test, or a predictor of job performance. The task's published validity comes from group-level correlations in research samples; individual scores are noisy, and scores move with mood, incentives, and even instructions. Treat any single run as one data point in a conversation, never as a verdict.
Two decades of popping balloons for data
Where the task came from, why behavioural measures beat questionnaires at their own game, and what the evidence actually shows.
In 2002, Carl Lejuez and colleagues wanted a measure of risk-taking that didn't rely on people describing themselves. Their answer was disarmingly simple: a balloon, a pump, and money that vanishes.
Why a balloon?
Questionnaires ask people what they would do; the Balloon Analogue Risk Task (BART) watches what they actually do when every choice carries real tension between growth and loss. The balloon metaphor works because it mirrors the structure of many real risks. Smoking one more cigarette, holding a position one more day, taking one more shortcut: each additional step pays a little more, until the step that costs everything.
In the original evaluation, BART scores correlated with self-reported smoking, drinking, drug use, gambling, unprotected sex and seat-belt neglect, over and above what standard questionnaires explained (Lejuez et al., 2002). That incremental validity is the task's claim to fame: it captures variance that self-report misses.
How the task is built
The standard protocol presents 30 balloons. Each pump pays 5¢ into a temporary reserve; the participant may bank the reserve at any time. Burst points are drawn uniformly from 1 to 128 pumps, so the average balloon can absorb 64 pumps, and the expected-value-maximising strategy is to pump each balloon into the 60s. Participants are told the balloon can pop at any point, but never the distribution: the uncertainty is the instrument.
The primary score is the adjusted average pumps: mean pumps on banked balloons only (see the score guide for why popped trials are excluded).
What the evidence shows
- Reliability. Test–retest correlations around .77, with practice effects small after the first session (White, Lejuez & de Wit, 2008).
- Convergent validity. A meta-analysis of 34 samples found reliable associations between BART scores and both sensation seeking and impulsivity, with the strongest links to real-world risk behaviour (Lauriola et al., 2014).
- Clinical sensitivity. Elevated scores in substance-using populations, adolescent risk-takers and antisocial profiles (Hunt et al., 2005).
- Neural correlates. Risk escalation on the BART engages mesolimbic reward circuitry; imaging work shows striatal and anterior-cingulate activation scaling with risk level (Rao et al., 2008).
Known limitations
The BART is one task, and honest use requires naming its edges. Scores conflate risk preference with learning speed: someone may pump little because they are cautious, or because they haven't yet worked out that early pops are cheap information. The task measures risk under escalating gains with total loss, one specific risk structure among many; a person bold with balloons may be timid with social risk, and vice versa. Laboratory stakes are small, and behaviour under 5¢ increments may not scale to consequential decisions. And like every behavioural task, single-session scores carry meaningful measurement noise.
None of this diminishes the task's research value. It does mean that any individual score (including yours) deserves a modest interpretation.
This implementation
This site follows the canonical parameters: uniform burst points over 1–128 pumps, 5¢ per pump, 30 balloons by default, with the standard instruction script adapted from the original paper. Configurable deviations (fewer balloons, different pay, reduced instructions) are available for demonstration purposes and are flagged as non-standard wherever scores are shown. Burst points are drawn with a cryptographically seeded PRNG at balloon start; nothing about the current balloon's burst point is knowable from the display. All data stays in your browser; nothing is uploaded, ever.
- Lejuez, C. W., Read, J. P., Kahler, C. W., Richards, J. B., Ramsey, S. E., Stuart, G. L., Strong, D. R., & Brown, R. A. (2002). Evaluation of a behavioral measure of risk taking: The Balloon Analogue Risk Task (BART). Journal of Experimental Psychology: Applied, 8(2), 75–84.
- White, T. L., Lejuez, C. W., & de Wit, H. (2008). Test–retest characteristics of the Balloon Analogue Risk Task (BART). Experimental and Clinical Psychopharmacology, 16(6), 565–570.
- Lauriola, M., Panno, A., Levin, I. P., & Lejuez, C. W. (2014). Individual differences in risky decision making: A meta-analysis of sensation seeking and impulsivity with the Balloon Analogue Risk Task. Journal of Behavioral Decision Making, 27(1), 20–36.
- Hunt, M. K., Hopko, D. R., Bare, R., Lejuez, C. W., & Robinson, E. V. (2005). Construct validity of the Balloon Analog Risk Task (BART): Associations with psychopathy and impulsivity. Assessment, 12(4), 416–428.
- Rao, H., Korczykowski, M., Pluta, J., Hoang, A., & Detre, J. A. (2008). Neural correlates of voluntary and involuntary risk taking in the human brain: An fMRI study of the Balloon Analog Risk Task (BART). NeuroImage, 42(2), 902–910.
Configure & share a session
Build a fixed-settings link, send it to a candidate, and read the run together afterwards. Every option below changes what the task measures, so choose deliberately.
Session settings
Settings are encoded in the link: nothing to install, no accounts.
Running it in an interview
- Before
- Ask for consent plainly: "this is a short decision game, it takes five minutes, and we'll look at the result together." Screen-share or send the link. Don't coach strategy.
- During
- Watch process, not just outcome: do they settle into a rhythm or decide pump-by-pump? What happens on the balloon after a pop: collapse to tiny banks, or a measured return?
- After
- The debrief is where the value is. "You banked at 12 pumps every time: what were you protecting?" or "You went back to 40 pumps straight after that pop; walk me through it." The reasoning is the signal; the score is just the prompt.
Ground rules
- Not a gate
- The BART has no validated cut-offs for employment. Never use a score to screen candidates in or out; use it to open a conversation about judgement under uncertainty.
- Comparability
- Compare candidates only on identical settings, and expect noise: single-session scores wobble. A second run tells you more than a decimal place ever will.
- Privacy
- Results live only in the candidate's browser. If you want the data, ask them to use the CSV export or copy the summary; it's theirs to share.