How Editorial Confidence Is Checked Against Candidate Evidence
Editorial confidence is not calibrated to observed question difficulty in the current bank. Well-sampled questions marked 100 confidence average 57.6% observed accuracy, compared with 58.2% below 50 confidence and 61% in the 50-74.9 band. Most active questions are marked 100.
What this result includes
- 513 questions at Below 50 confidence
- 845 questions at 50-74.9 confidence
- 4,809 questions at 100 confidence
The field is heavily concentrated at 100
4,809 of 6,167 active IDs sit in that band.
Confidence is not difficulty
A well-sourced question can still be difficult. The current values should trigger source, wording, distractor, and mapping review rather than a score claim.
A multidimensional rubric would be more useful
Source certainty, wording quality, answer uniqueness, and objective mapping should be recorded separately before this field becomes a public validation measure.
Editorial confidence and observed accuracy
Observed accuracy uses only questions with at least 30 active learners and 100 attempts.
Confidence calibration3 rows
Editorial confidence describes source and authoring certainty. Candidate accuracy measures difficulty; neither should substitute for an item review.
| Confidence | Active questions | Well sampled | Avg. confidence | Observed accuracy |
|---|---|---|---|---|
| Below 50 | 513 | 114 | 25.5% | 58.2% |
| 50-74.9 | 845 | 166 | 66.7% | 61.0% |
| 100 | 4,809 | 358 | 100.0% | 57.6% |
What the data cannot establish
Most active questions currently use confidence 100, so the field is not discriminating enough to publish as a validated score.
Turn the finding into a decision.
Use the report to choose a focused action, then test that decision with current material. The evidence should reduce uncertainty; it should never replace the complete exam outline or promise an official result.