Skip to main content

How it works · v4.3.2

How we measure exam readiness

“Percent correct” is easy to misread — it depends on which questions you happened to get. Our readiness index answers a harder, fairer question: at your current ability, weighted like the real exam, how ready are you — and how sure are we? It rests on measurement theory (IRT), Bayesian ability estimation, and honest uncertainty — the same scientific family used to build real board exams.

In one line: The headline is retained exam-day ability (τ_ret): expected accuracy on the published question pools at your current ability, including only earned learning bonuses. Cold first-try ability (τ_cold) appears as a secondary only when its rounded display differs. R = A × τ_ret is the full-board lower bound; when earned A cannot reach the benchmark, the verdict is Bank-limited. This is not a pass probability.

Cold vs retained

τ_cold records the first-try baseline. τ_ret is the headline because delayed, earned learning can improve exam-day expectations; the cold value appears only when the two rounded displays differ.

What each number means

  • θ · Latent ability

    EAP ability per BCSC volume from first-try answers — how hard items you can handle cold.

  • τ_cold · Cold first-try ability

    Published-pool expected accuracy inferred from first tries — secondary, and shown only when its rounded display differs from τ_ret.

  • τ_ret · Retained / exam-day ability

    Headline score: published-pool accuracy at θ plus earned learning bonuses; per-item FSRS retrievability ρ scales only those bonuses.

  • A · Earned bank coverage

    Graded share of blueprint weight the published pools can fill; each volume earns weight in proportion to the share of its board draw its pool could supply.

  • R · Full-board lower bound

    R = A × τ_ret. Uncovered blueprint weight counts as 0 — the bank ceiling on any board claim.

  • F · Aggregate freshness

    Diagnostic recall summary across practised, answerable content. F is reported separately and does not scale τ_ret or R.

  • π · Pass probability

    Calibrated P(pass) for a named exam — only after external outcomes validate a proper scoring-rule model. Until then: null.

How the parts combine

Every symbol has a fixed place in the algebra. We combine them only with named formulas — never an opaque mashup like “0.4·ability + 0.3·freshness”. The headline is always τ_ret, τ_cold is conditional secondary context, and full-board claims use R = A × τ_ret.

  1. 1. θ_dEAP(first-try | b, c) · Latent ability per volume
  2. 2. p̄_dE[correct | θ̂_d] on published pool · Cold volume expected accuracy
  3. 3. τ_coldΣ(w·r·p̄) / Σ(w·r), mock-blend · Cold first-try ability (secondary)
  4. 4. learned≥2 delayed corrects after miss + (S≥3d or spaced pair) · Anti-gaming learning evidence
  5. 5. ρ_iFSRS R(t,S) to exam day · Exam-day retrievability
  6. 6. τ_retΣ(w·q·r·mean[p̄ + ρ·max(0, floor_learned − p̄)]) / Σ(w·q·r) · Retained / exam-day ability (headline)
  7. 7. AΣ min(w_d, pool_d / blueprint_total) / Σ w_d · Earned bank ceiling
  8. 8. RA × τ_ret · Full-board lower bound
  9. 9. CI(R)A · logitN(τ_ret, SE_ci) ∩ [0, A] · Uncertainty
  10. 10. bandside of pass line the CI clears (headline and each volume) · Claim only when interval allows
  11. 11. πcalibrated P(pass | features) · Null until holdout gate
From answers to readiness
Your first-try answersCalibrate eachquestion's difficultyEstimate your abilityper BCSC volumeScore volumes at yourability levelWeight by the examblueprintCredit proven learning +FSRS ρτ_ret headline,conditional τ_cold, R =…

The pipeline, end to end. Each stage is explained below.

1

Why raw “% correct” misleads

Score 80% on an easy set and 55% on a brutal one and you might be equally ready — the number moved because the questions changed, not you. We also count only your first attempt on each question (not Review re-tries), because that mirrors a cold exam encounter. So a plain average is the wrong unit. We need to measure you, independent of the paper you drew.

2

Every question gets a fair difficulty

Each question starts from the author's Easy/Medium/Hard rating, then we shrink that toward what the whole cohort actually scores (empirical-Bayes, on a Rasch scale). Popular questions are trusted more; new or rarely-answered ones lean on the author rating and are flagged preliminary. This is the same shrinkage idea (Stein / empirical Bayes) that keeps small-sample estimates from overreacting.

3

We estimate your ability, not your luck

With calibrated difficulties, Item Response Theory backs out your latent ability (θ) for each BCSC volume — separating how hard the questions were from how well you did. A guessing floor (1 ÷ number of options) stops multiple-choice luck from inflating θ. Every first attempt counts equally, whatever you did with the question afterwards: weighting this evidence by review activity would quietly favour the questions you got right, because those are the ones you never sent back to review.

4

Scored like the real exam

Each volume remains a published-pool expectation at its estimated ability. Volume expectations are combined using graded, earned bank weight: a pool earns only the share of its blueprint draw that it could fill. That earned share is A, and unfillable blueprint weight contributes 0 to R = A × τ_ret.

5

Memory can only add, never subtract

First tries estimate cold ability. A missed item can earn a learning bonus only through delayed spaced success; FSRS exam-day retrievability (ρ) scales that bonus alone. It does not multiply cold ability. As the bonus fades, τ_ret returns to τ_cold, never below it.

Aggregate freshness F summarizes recall across practised, answerable content for diagnosis and study planning. It is reported separately and is not another multiplier in τ_ret or R; only per-item ρ scales a proven-learning bonus.

6

An honest number, with an interval

The headline carries an approximate 95% logit-normal interval, widened for preliminary item calibration; multiplying its endpoints by A gives the interval on R. This is a study compass, not a calibrated probability of passing. The estimate unlocks only after: ~30 unique first tries, at least 35% of blueprint weight earned by the bank, ≥50% sampling of that answerable bank across ≥3 volumes, and a tight enough ability SE — so grinding one topic or a thin bank can't manufacture a confident board score. A recent full-length mock may waive sampling and breadth, but not the question count, bank floor, or precision.

Each volume row shows retained expected accuracy on that volume’s published pool. Its own delta-method interval—not the point estimate—places it Below goal, Too close to call, or At/above goal. Thin pools stay approximate as our coverage constraint; when aggregate A is below the goal, the headline is Bank-limited—not a learner result.

The honesty gate
New readiness estimateQs + bank +coverage/breadth +SE?Show progress, keepscore lockedReveal retained abilityτ_retNot yetYes

Put together: calibrated questions → latent ability θ → pool expectations → τ_ret headline → graded earned A → R = A × τ_ret → interval-based claims. τ_cold appears only when its rounded display differs; F remains diagnostic.

Scientific axioms

  1. Latent trait: competence is θ; answers are noisy measurements via a 3PL-style response model, not raw % correct.
  2. Bayes: θ̂ is EAP with a weakly informative prior; sparse volumes borrow the global mean but not other volumes' precision.
  3. Unbiased evidence: every first try enters the likelihood with equal weight. Weighting evidence by anything correlated with being right or wrong (such as whether an item has a review card) would bias ability upward.
  4. Uncertainty: every unlocked headline carries a SE/CI; mock↔practice blends keep a Cov floor so intervals cannot look falsely sharp.
  5. Separability: ability τ, bank coverage A, sampling C, freshness F, and pass probability π are distinct estimands — never mashed into one opaque score.
  6. Lower bound: R = A × τ; blueprint weight we cannot fill with published items contributes 0 to full-board readiness.
  7. Earned weight: a volume claims blueprint weight only in proportion to the share of a real board draw its published pool could fill — five items never stand in for 14% of the exam.
  8. First-try only: practice first attempts enter the likelihood (exam-like cold measurement).
  9. Memory adds, never subtracts: proving you learned a missed item can only raise its expected accuracy toward a ceiling, and forgetting returns you to your cold baseline.
  10. No uncalibrated π: pass_likelihood stays null until external outcomes validate a proper scoring-rule model.
  11. Anti-gaming: breadth, blueprint weights, and drill priority priced on answerable weight (not nominal weight) resist one-volume cherry-picking.
  12. No claim without an interval, at every level: a volume is called above or below its goal only when that volume's own interval clears it, and a pool too thin to fill its board slot is labelled as our coverage limit rather than the user's result.

This is one half of the system. The other is adaptive spaced repetition — how we decide what to show you and when. Both rest on clinician-written, reviewed questions. See it all on the how-it-works overview.

The science behind it

The index rests on established psychometrics — the same family of methods used to build and score real standardized exams: