Case study · 2026 · AI · Web · Custom software
A wrong multiple-choice answer tells you nothing about the reasoning behind it.Built a radar that reads open-ended answers and tells the teacher which misconceptions are running the room.
ClarixAI processes open-ended student answers in Dutch, English and Hungarian through a per-language transformer ensemble (RobBERT-2023, RoBERTa-base, huBERT) with a 1-vs-rest sigmoid head over eight misconception meta-tags. HDBSCAN clustering on encoder embeddings surfaces structural patterns across questions; the dashboard turns each tag into a what-it-means / why-it-happens / what-to-do card. The studio shipped the training pipeline, the inference service, and the teacher dashboard.
- Python
- PyTorch
- Transformers
- FastAPI
- React

Inside the build



Overview
- 3
- Languages · NL EN HU
- 8
- Misconception meta-tags
- 4-ckpt
- Ensemble · NL prod
- HDBSCAN
- Cross-question clustering
ClarixAI reads open-ended student answers (Dutch, English, Hungarian) and surfaces the misconception patterns dominating a class - concept confusion, surface-pattern matching, missing precondition, and six others - with a concrete pedagogical action card per pattern. It does not grade students; it grades the class's reasoning. A per-language multi-task transformer ensemble plus HDBSCAN clustering finds structural patterns across questions.
What shipped
What it does
- Multi-task misconception classifier · NL · EN · HU
- 8 reasoning-error meta-tags · with a teacher action per tag
- HDBSCAN clustering for cross-question structural patterns
- Cohort-level trend dashboard
- Pedagogy library · what it means, why it happens, what to do
The problem
- Multiple-choice tells you 'wrong' · not WHY they were wrong
- Open-ended answers are too time-consuming to read for patterns
- 'More practice' is the default action when teachers can't see the pattern
Why it matters
- The class's dominant reasoning errors are visible per question
- Every error pattern comes with a concrete teacher response
- Cohort trends make curriculum gaps obvious
How it shipped
- 01 · BRIEF
Pin the unit of analysis as the reasoning behind the answer - not the answer itself.
Decision the product rests on: the system never grades students. It surfaces the misconceptions running through the class so the teacher knows which question to re-explain, and how.
- 02 · BUILD
Per-language transformer + anti-overfit recipe + ensemble for prod.
RobBERT-2023 for Dutch, RoBERTa-base for English, huBERT for Hungarian - trained with the v5.0 anti-overfit recipe (frozen lower layers, focal loss, SupCon, weighted gold). NL ships as a 4-checkpoint ensemble (v6.2) with bounded per-tag thresholds + a k-NN prototype blend; EN matches Dutch on gold F1 and beats it on adversarial argmax.
- 03 · SHIP
FastAPI inference + Vite/React dashboard, three languages live.
Inference clusters per-question answers via HDBSCAN over the encoder's embeddings, surfaces structural cross-question patterns and per-cohort trends, and feeds the pedagogical knowledge base that turns each tag into a teacher action card.
Stack
Model
Per-language ensemble · 8 misconception tags
RobBERT / RoBERTa / huBERT with a 1-vs-rest sigmoid head - concept confusion, surface pattern, missing precondition + five more.
Pipeline
Classify → cluster → trend
Classify each answer, HDBSCAN-cluster the encoder embeddings per question, then roll up the cross-question and cohort-level patterns.
Dashboard
Per-question + cohort trend
Vite/React with i18n (nl/hu/en) - teachers see misconception density per question and the trend across cohorts side by side.
Pedagogy
Action card per misconception tag
Every tag links to a what-it-means · why-it-happens · what-to-do card - the radar surfaces the pattern, the card tells the teacher how to respond.
Case study
“Our open-ended questions used to be the part of the assessment we'd skim, because reading 60 answers for a pattern was a whole evening. The team built a system that does the reading for us and tells us, per question, which misconception is running the room - and what to do about it, not just that it's there. We stopped over-prescribing 'more practice'. We started reteaching the right thing.”
More work
Next project
GlowUp