How it works · v4.3.2
How we measure exam readiness
“Percent correct” is easy to misread — it depends on which questions you happened to get. Our readiness index answers a harder, fairer question: at your current ability, weighted like the real exam, how ready are you — and how sure are we? It rests on measurement theory (IRT), Bayesian ability estimation, and honest uncertainty — the same scientific family used to build real board exams.
In one line: The headline is retained exam-day ability (τ_ret): expected accuracy on the published question pools at your current ability, including only earned learning bonuses. Cold first-try ability (τ_cold) appears as a secondary only when its rounded display differs. R = A × τ_ret is the full-board lower bound; when earned A cannot reach the benchmark, the verdict is Bank-limited. This is not a pass probability.
Cold vs retained
τ_cold records the first-try baseline. τ_ret is the headline because delayed, earned learning can improve exam-day expectations; the cold value appears only when the two rounded displays differ.
What each number means
θ · Latent ability
EAP ability per BCSC volume from first-try answers — how hard items you can handle cold.
τ_cold · Cold first-try ability
Published-pool expected accuracy inferred from first tries — secondary, and shown only when its rounded display differs from τ_ret.
τ_ret · Retained / exam-day ability
Headline score: published-pool accuracy at θ plus earned learning bonuses; per-item FSRS retrievability ρ scales only those bonuses.
A · Earned bank coverage
Graded share of blueprint weight the published pools can fill; each volume earns weight in proportion to the share of its board draw its pool could supply.
R · Full-board lower bound
R = A × τ_ret. Uncovered blueprint weight counts as 0 — the bank ceiling on any board claim.
F · Aggregate freshness
Diagnostic recall summary across practised, answerable content. F is reported separately and does not scale τ_ret or R.
π · Pass probability
Calibrated P(pass) for a named exam — only after external outcomes validate a proper scoring-rule model. Until then: null.
How the parts combine
Every symbol has a fixed place in the algebra. We combine them only with named formulas — never an opaque mashup like “0.4·ability + 0.3·freshness”. The headline is always τ_ret, τ_cold is conditional secondary context, and full-board claims use R = A × τ_ret.
- 1. θ_d — EAP(first-try | b, c) · Latent ability per volume
- 2. p̄_d — E[correct | θ̂_d] on published pool · Cold volume expected accuracy
- 3. τ_cold — Σ(w·r·p̄) / Σ(w·r), mock-blend · Cold first-try ability (secondary)
- 4. learned — ≥2 delayed corrects after miss + (S≥3d or spaced pair) · Anti-gaming learning evidence
- 5. ρ_i — FSRS R(t,S) to exam day · Exam-day retrievability
- 6. τ_ret — Σ(w·q·r·mean[p̄ + ρ·max(0, floor_learned − p̄)]) / Σ(w·q·r) · Retained / exam-day ability (headline)
- 7. A — Σ min(w_d, pool_d / blueprint_total) / Σ w_d · Earned bank ceiling
- 8. R — A × τ_ret · Full-board lower bound
- 9. CI(R) — A · logitN(τ_ret, SE_ci) ∩ [0, A] · Uncertainty
- 10. band — side of pass line the CI clears (headline and each volume) · Claim only when interval allows
- 11. π — calibrated P(pass | features) · Null until holdout gate
The pipeline, end to end. Each stage is explained below.
Why raw “% correct” misleads
Score 80% on an easy set and 55% on a brutal one and you might be equally ready — the number moved because the questions changed, not you. We also count only your first attempt on each question (not Review re-tries), because that mirrors a cold exam encounter. So a plain average is the wrong unit. We need to measure you, independent of the paper you drew.
Every question gets a fair difficulty
Each question starts from the author's Easy/Medium/Hard rating, then we shrink that toward what the whole cohort actually scores (empirical-Bayes, on a Rasch scale). Popular questions are trusted more; new or rarely-answered ones lean on the author rating and are flagged preliminary. This is the same shrinkage idea (Stein / empirical Bayes) that keeps small-sample estimates from overreacting.
We estimate your ability, not your luck
With calibrated difficulties, Item Response Theory backs out your latent ability (θ) for each BCSC volume — separating how hard the questions were from how well you did. A guessing floor (1 ÷ number of options) stops multiple-choice luck from inflating θ. Every first attempt counts equally, whatever you did with the question afterwards: weighting this evidence by review activity would quietly favour the questions you got right, because those are the ones you never sent back to review.
Scored like the real exam
Each volume remains a published-pool expectation at its estimated ability. Volume expectations are combined using graded, earned bank weight: a pool earns only the share of its blueprint draw that it could fill. That earned share is A, and unfillable blueprint weight contributes 0 to R = A × τ_ret.
Memory can only add, never subtract
First tries estimate cold ability. A missed item can earn a learning bonus only through delayed spaced success; FSRS exam-day retrievability (ρ) scales that bonus alone. It does not multiply cold ability. As the bonus fades, τ_ret returns to τ_cold, never below it.
Aggregate freshness F summarizes recall across practised, answerable content for diagnosis and study planning. It is reported separately and is not another multiplier in τ_ret or R; only per-item ρ scales a proven-learning bonus.
An honest number, with an interval
The headline carries an approximate 95% logit-normal interval, widened for preliminary item calibration; multiplying its endpoints by A gives the interval on R. This is a study compass, not a calibrated probability of passing. The estimate unlocks only after: ~30 unique first tries, at least 35% of blueprint weight earned by the bank, ≥50% sampling of that answerable bank across ≥3 volumes, and a tight enough ability SE — so grinding one topic or a thin bank can't manufacture a confident board score. A recent full-length mock may waive sampling and breadth, but not the question count, bank floor, or precision.
Each volume row shows retained expected accuracy on that volume’s published pool. Its own delta-method interval—not the point estimate—places it Below goal, Too close to call, or At/above goal. Thin pools stay approximate as our coverage constraint; when aggregate A is below the goal, the headline is Bank-limited—not a learner result.
Put together: calibrated questions → latent ability θ → pool expectations → τ_ret headline → graded earned A → R = A × τ_ret → interval-based claims. τ_cold appears only when its rounded display differs; F remains diagnostic.
Scientific axioms
- Latent trait: competence is θ; answers are noisy measurements via a 3PL-style response model, not raw % correct.
- Bayes: θ̂ is EAP with a weakly informative prior; sparse volumes borrow the global mean but not other volumes' precision.
- Unbiased evidence: every first try enters the likelihood with equal weight. Weighting evidence by anything correlated with being right or wrong (such as whether an item has a review card) would bias ability upward.
- Uncertainty: every unlocked headline carries a SE/CI; mock↔practice blends keep a Cov floor so intervals cannot look falsely sharp.
- Separability: ability τ, bank coverage A, sampling C, freshness F, and pass probability π are distinct estimands — never mashed into one opaque score.
- Lower bound: R = A × τ; blueprint weight we cannot fill with published items contributes 0 to full-board readiness.
- Earned weight: a volume claims blueprint weight only in proportion to the share of a real board draw its published pool could fill — five items never stand in for 14% of the exam.
- First-try only: practice first attempts enter the likelihood (exam-like cold measurement).
- Memory adds, never subtracts: proving you learned a missed item can only raise its expected accuracy toward a ceiling, and forgetting returns you to your cold baseline.
- No uncalibrated π: pass_likelihood stays null until external outcomes validate a proper scoring-rule model.
- Anti-gaming: breadth, blueprint weights, and drill priority priced on answerable weight (not nominal weight) resist one-volume cherry-picking.
- No claim without an interval, at every level: a volume is called above or below its goal only when that volume's own interval clears it, and a pool too thin to fill its board slot is labelled as our coverage limit rather than the user's result.
This is one half of the system. The other is adaptive spaced repetition — how we decide what to show you and when. Both rest on clinician-written, reviewed questions. See it all on the how-it-works overview.
The science behind it
The index rests on established psychometrics — the same family of methods used to build and score real standardized exams:
- Probabilistic Models for Some Intelligence and Attainment Tests
Rasch, G. (1960) — The Rasch model — separates item difficulty from person ability.
- Applications of Item Response Theory to Practical Testing Problems
Lord, F. M. (1980) — Foundational text on Item Response Theory (IRT) for exams.
- Item Response Theory for Psychologists
Embretson, S. E., & Reise, S. P. (2000) — Accessible treatment of ability estimation and item calibration.
- Adaptive EAP estimation of ability in a microcomputer environment
Bock, R. D., & Mislevy, R. J. (1982) — Applied Psychological Measurement — the EAP posterior behind our interval.
- Stein's paradox in statistics
Efron, B., & Morris, C. (1977) — Scientific American — the empirical-Bayes shrinkage we use for difficulty.
- Logistic-normal distributions: Some properties and uses
Aitchison, J., & Shen, S. M. (1980) — Biometrika — the transformed-normal family used for the headline interval.
- A note on the delta method
Oehlert, G. W. (1992) — The American Statistician — the approximation used for per-volume uncertainty.
- A Stochastic Shortest Path Algorithm for Optimizing Spaced Repetition Scheduling
Ye, J., et al. (2022) — FSRS — the power-law retrievability model that scales proven-learning bonuses.