DFIELDSOLUTIONS

Case study · 2026 · AI · Web · Custom software

A wrong multiple-choice answer tells you nothing about the reasoning behind it.Built a radar that reads open-ended answers and tells the teacher which misconceptions are running the room.

ClarixAI processes open-ended student answers in Dutch, English and Hungarian through a per-language transformer ensemble (RobBERT-2023, RoBERTa-base, huBERT) with a 1-vs-rest sigmoid head over eight misconception meta-tags. HDBSCAN clustering on encoder embeddings surfaces structural patterns across questions; the dashboard turns each tag into a what-it-means / why-it-happens / what-to-do card. The studio shipped the training pipeline, the inference service, and the teacher dashboard.

  • Python
  • PyTorch
  • Transformers
  • FastAPI
  • React
Internal build, no public URL
ClarixAI

Inside the build

ClarixAI — Opening state
Opening state
ClarixAI — In use
In use
ClarixAI — Result
Result

Overview

3
Languages · NL EN HU
8
Misconception meta-tags
4-ckpt
Ensemble · NL prod
HDBSCAN
Cross-question clustering

ClarixAI reads open-ended student answers (Dutch, English, Hungarian) and surfaces the misconception patterns dominating a class - concept confusion, surface-pattern matching, missing precondition, and six others - with a concrete pedagogical action card per pattern. It does not grade students; it grades the class's reasoning. A per-language multi-task transformer ensemble plus HDBSCAN clustering finds structural patterns across questions.

What shipped

What it does

  • Multi-task misconception classifier · NL · EN · HU
  • 8 reasoning-error meta-tags · with a teacher action per tag
  • HDBSCAN clustering for cross-question structural patterns
  • Cohort-level trend dashboard
  • Pedagogy library · what it means, why it happens, what to do

The problem

  • Multiple-choice tells you 'wrong' · not WHY they were wrong
  • Open-ended answers are too time-consuming to read for patterns
  • 'More practice' is the default action when teachers can't see the pattern

Why it matters

  • The class's dominant reasoning errors are visible per question
  • Every error pattern comes with a concrete teacher response
  • Cohort trends make curriculum gaps obvious

How it shipped

  1. 01 · BRIEF

    Pin the unit of analysis as the reasoning behind the answer - not the answer itself.

    Decision the product rests on: the system never grades students. It surfaces the misconceptions running through the class so the teacher knows which question to re-explain, and how.

  2. 02 · BUILD

    Per-language transformer + anti-overfit recipe + ensemble for prod.

    RobBERT-2023 for Dutch, RoBERTa-base for English, huBERT for Hungarian - trained with the v5.0 anti-overfit recipe (frozen lower layers, focal loss, SupCon, weighted gold). NL ships as a 4-checkpoint ensemble (v6.2) with bounded per-tag thresholds + a k-NN prototype blend; EN matches Dutch on gold F1 and beats it on adversarial argmax.

  3. 03 · SHIP

    FastAPI inference + Vite/React dashboard, three languages live.

    Inference clusters per-question answers via HDBSCAN over the encoder's embeddings, surfaces structural cross-question patterns and per-cohort trends, and feeds the pedagogical knowledge base that turns each tag into a teacher action card.

Stack

Model

Per-language ensemble · 8 misconception tags

RobBERT / RoBERTa / huBERT with a 1-vs-rest sigmoid head - concept confusion, surface pattern, missing precondition + five more.

Pipeline

Classify → cluster → trend

Classify each answer, HDBSCAN-cluster the encoder embeddings per question, then roll up the cross-question and cohort-level patterns.

Dashboard

Per-question + cohort trend

Vite/React with i18n (nl/hu/en) - teachers see misconception density per question and the trend across cohorts side by side.

Pedagogy

Action card per misconception tag

Every tag links to a what-it-means · why-it-happens · what-to-do card - the radar surfaces the pattern, the card tells the teacher how to respond.

Case study

“Our open-ended questions used to be the part of the assessment we'd skim, because reading 60 answers for a pattern was a whole evening. The team built a system that does the reading for us and tells us, per question, which misconception is running the room - and what to do about it, not just that it's there. We stopped over-prescribing 'more practice'. We started reteaching the right thing.”

Anonymous · Researcher · education-AI platform (under NDA) · NL

More work

DField Solutions · DField Bt. · dezso@dfieldsolutions.com
5.0
“From LinkedIn DM to live site. Two tiny tweaks, then shipped.”Michael J Ringer · Vilya ProtectionFounder · Spain