How the score is built
Most scoring products return one number and no reasoning. Here is exactly what happens to a response you submit — and what we do when we are not sure.
| # | Layer | What it does |
|---|---|---|
| 1 | Deterministic checks | Word counts, structural checks, and the triggers that score an automatic zero. No AI is involved and nothing here can be wrong in an interesting way. |
| 2 | Measurement | Transcription, acoustic pronunciation and fluency measurement, and deterministic grammar and agreement checking. |
| 3 | Evidence | Every measurement is assembled into one structured record and handed to the judges as evidence, not as instructions. |
| 4 | Two independent judges | Two separate models score against the official criteria plus calibration anchors. Neither can see the other's score. |
| 5 | Disagreement | One level apart, we combine and narrow. Two apart, we widen the range and say the result is uncertain. Three or more, no number is shown until a third judge rules, and you are offered a free human review. |
| 6 | Calibration | Scores are corrected against a curve fitted to real graded samples, including a per-language-group correction where there is enough data to measure one. |
| 7 | The range | The output is an interval with a coverage guarantee, not a point estimate. |
We never present the average of two very different scores as if it were settled. That is the single most common way a scoring product misleads someone.
What actually produces the number
No human examiner reads your work here. Every score on this product is produced by software, and the honest way to describe that is to say which piece does which job.
Running today
| Job | How it is done |
|---|---|
| Turning your recording into text | Automatic speech recognition |
| Pronunciation, stress, fluency and hesitation | Acoustic scoring at the phoneme level — per sound, not per sentence |
| Grammar, spelling and agreement | Deterministic rule checking. No model is involved, so it cannot be confidently wrong |
| Applying the exam's published criteria to your writing and speaking | A language model, given the deterministic measurements above as evidence rather than as instructions |
The processors that carry out this work are named, with what each one receives and why, on the data and processors page inside the app. That page is kept current with the stack; this one deliberately describes the function rather than the supplier, so it cannot quietly go out of date.
Being built
- Two independent language models score every response separately. Neither can see the other's score. Today one model does the scoring; the second judge and the arbitration below are the part under construction.
- When the two disagree by one level, the scores are combined and the range narrows.
- When they disagree by two levels, the range widens and the product says so rather than presenting an average as if it were settled.
- When they disagree by three levels or more, no number is shown at all until a third model arbitrates — and a free human review is offered.
- Calibration and the confidence interval. Raw scores corrected against a curve fitted to real graded samples, then converted into an interval with a coverage guarantee.
What the market does and does not publish:
France Éducation international, through TV5MONDE, publishes more than 600 TCF practice questions free, with an offline simulator. Commercial platforms exist, and several score written production automatically. Not one publishes a measured accuracy figure — no corpus, no rater protocol, no agreement statistic. The market competes on outcome claims about its candidates. We intend to compete on published, measured error:
We will measure this system's error against real official score reports and publish it — including when the result does not flatter us.
Written in the future tense on purpose. It requires a calibration set of real score reports that does not exist yet, and the first publication comes at the close of the French phase. Claiming it in the present tense today would be exactly the kind of unearned accuracy claim this product exists to replace.
What we promise, and what we do not
We do not claim to predict your official result. We never say “guaranteed”, and we never claim to be 100 % accurate — no one honestly can, and a product that says so is telling you something about itself.
In build What we will publish at the close of the French phase: our mean absolute error against real official score reports, the agreement statistic between our scores and human raters, and how often our stated 90 % range actually contained the real result. Plus a fairness audit by candidate first language, including the groups where we perform worse.
If we do not hit our own accuracy thresholds, we will not publish a numeric prediction at all. The product gives qualitative feedback only until we do. That rule is written into our development plan as a release gate, not as an aspiration.
One limitation, stated plainly
The coverage guarantee on our range holds when the people we calibrated against resemble the people using the product. Our calibration group is recruited, so that match is close but not perfect. We re-estimate coverage against real user results every quarter, and we will publish it when it moves.