Free

Sit a free mock exam — no account needed. Real timing, one play of each recording, an exact comprehension score. Make an account at the end to read it. No card, no trial. IELTS and TCF Canada are built; TEF Canada is not started. Start the mock exam →

SELMKNOW YOUR SCORE

How the score is built

Most scoring products return one number and no reasoning. Here is exactly what happens to a response you submit — and what we do when we are not sure.

HOW ONE RESPONSE BECOMES ONE RANGE Outlined steps are deterministic · the filled step is judged · amber is the disagreement branch 1 Deterministic checks 2 Measurement 3 Evidence package 4 Two independent judges 5 Disagreement handling 6 Calibration 7 Conformal interval WHEN THE JUDGES DISAGREE 1 level apart → combine and narrow 2 apart → widen, and say so 3 or more → no number at all a third judge, and a free review
Layers 1–3 are deterministic and cost nothing. Layer 4 is where the money goes. Layers 5–7 are what make the output honest.
#LayerWhat it does
1Deterministic checksWord counts, structural checks, and the triggers that score an automatic zero. No AI is involved and nothing here can be wrong in an interesting way.
2MeasurementTranscription, acoustic pronunciation and fluency measurement, and deterministic grammar and agreement checking.
3EvidenceEvery measurement is assembled into one structured record and handed to the judges as evidence, not as instructions.
4Two independent judgesTwo separate models score against the official criteria plus calibration anchors. Neither can see the other's score.
5DisagreementOne level apart, we combine and narrow. Two apart, we widen the range and say the result is uncertain. Three or more, no number is shown until a third judge rules, and you are offered a free human review.
6CalibrationScores are corrected against a curve fitted to real graded samples, including a per-language-group correction where there is enough data to measure one.
7The rangeThe output is an interval with a coverage guarantee, not a point estimate.

We never present the average of two very different scores as if it were settled. That is the single most common way a scoring product misleads someone.

What actually produces the number

No human examiner reads your work here. Every score on this product is produced by software, and the honest way to describe that is to say which piece does which job.

Running today

JobHow it is done
Turning your recording into textAutomatic speech recognition
Pronunciation, stress, fluency and hesitationAcoustic scoring at the phoneme level — per sound, not per sentence
Grammar, spelling and agreementDeterministic rule checking. No model is involved, so it cannot be confidently wrong
Applying the exam's published criteria to your writing and speakingA language model, given the deterministic measurements above as evidence rather than as instructions

The processors that carry out this work are named, with what each one receives and why, on the data and processors page inside the app. That page is kept current with the stack; this one deliberately describes the function rather than the supplier, so it cannot quietly go out of date.

Being built

  • Two independent language models score every response separately. Neither can see the other's score. Today one model does the scoring; the second judge and the arbitration below are the part under construction.
  • When the two disagree by one level, the scores are combined and the range narrows.
  • When they disagree by two levels, the range widens and the product says so rather than presenting an average as if it were settled.
  • When they disagree by three levels or more, no number is shown at all until a third model arbitrates — and a free human review is offered.
  • Calibration and the confidence interval. Raw scores corrected against a curve fitted to real graded samples, then converted into an interval with a coverage guarantee.

What the market does and does not publish:

France Éducation international, through TV5MONDE, publishes more than 600 TCF practice questions free, with an offline simulator. Commercial platforms exist, and several score written production automatically. Not one publishes a measured accuracy figure — no corpus, no rater protocol, no agreement statistic. The market competes on outcome claims about its candidates. We intend to compete on published, measured error:

We will measure this system's error against real official score reports and publish it — including when the result does not flatter us.

Written in the future tense on purpose. It requires a calibration set of real score reports that does not exist yet, and the first publication comes at the close of the French phase. Claiming it in the present tense today would be exactly the kind of unearned accuracy claim this product exists to replace.

What we promise, and what we do not

We do not claim to predict your official result. We never say “guaranteed”, and we never claim to be 100 % accurate — no one honestly can, and a product that says so is telling you something about itself.

In build What we will publish at the close of the French phase: our mean absolute error against real official score reports, the agreement statistic between our scores and human raters, and how often our stated 90 % range actually contained the real result. Plus a fairness audit by candidate first language, including the groups where we perform worse.

If we do not hit our own accuracy thresholds, we will not publish a numeric prediction at all. The product gives qualitative feedback only until we do. That rule is written into our development plan as a release gate, not as an aspiration.

One limitation, stated plainly

The coverage guarantee on our range holds when the people we calibrated against resemble the people using the product. Our calibration group is recruited, so that match is close but not perfect. We re-estimate coverage against real user results every quarter, and we will publish it when it moves.