Release summary
Specialized for the work of teaching
Teach-1.0 is optimized for diagnosing misconceptions, deciding when to scaffold versus when to hold back, sustaining multi-turn sessions, and producing measurable student-progress signals.
It is reported to be strongest on K–12 and early higher-education mathematics and STEM tutoring. Quantized deployment is intended to let schools, universities, and lower-resource institutions self-host without continuous API costs or student conversations leaving institutional infrastructure.
The model is derived from a strong open 4B-class base in the Qwen3-4B family and receives additional supervised and reinforcement-learning stages focused on pedagogical behavior.
Model card
Model details
Known release metadata is stated directly; fields not established by the supplied draft remain labeled as such.
Developer
AFAI Open Models
Release
Teach-1.0
Release date
August 25, 2026
Model type
Specialized pedagogical language model
Model size
4B class
Base model
Qwen3-4B family
Modality
Text
Primary domain
K–12 and early higher-education mathematics and STEM
License
Not specified in current release draft
Languages
Not specified in current release draft
Context length
Not disclosed
Release status
Research release / documentation in progress
Evaluation
Pedagogical performance
MainScore is described as a weighted aggregate of process quality, scaffolding decisions, misconception diagnosis, and simulated student-outcome proxies. Constraint-violating solutions receive zero.
Use guidance
Intended and out-of-scope uses
Intended uses
- Tutoring assistance with educator or institutional oversight
- Mathematics and STEM learning support
- Research on pedagogical dialogue and scaffolding
- Locally hosted educational prototypes
- Evaluation and improvement of tutoring workflows
Out of scope without further validation
- Replacing qualified teachers or human support
- High-stakes grading, admissions, or disciplinary decisions
- Safety-critical, medical, legal, or crisis guidance
- Unsupervised deployment to children
- Subjects and languages not adequately evaluated
- Assuming generated explanations are accurate
Development
How Teach-1.0 was trained
Stable long-horizon RL
Training techniques preserve policy entropy across longer rollouts and reduce train–inference mismatch, with the goal of avoiding repetitive or overly verbose behavior.
Pedagogical data pipeline
Trajectories receive automated and sampled human checks for verifier quality, difficulty calibration, and answer-dumping or reward exploitation.
Self-compaction
Near context limits, the model produces a concise working-state summary of student understanding, open goals, and recent misconceptions before continuing.
Efficiency pressure
Alternating phases optimize pedagogical success and then apply soft pressure on tokens, turns, and unnecessary elaboration.
Data
Training data & verification
The supplied draft reports a mixture of real de-identified tutoring transcripts and filtered synthetic trajectories generated under pedagogical constraints.
Reported quality controls
- Verifier false-positive and false-negative checks
- Difficulty calibration against the base model
- Answer-dumping and exploit detection
- Network isolation and grading-path separation
- Sampled human review
Documentation still needed
- Dataset names, sizes, and mixture proportions
- Transcript consent and de-identification process
- Collection geography and learner demographics
- Copyright, licensing, and retention details
- Bias and representation analysis
Observed behavior
Reported behavioral shifts
Explores student reasoning before intervening
Uses more condensed reasoning on routine steps
Probes assumptions and possible misconceptions
Asks targeted questions when student intent is ambiguous
Maintains more consistent help-versus-hold-back decisions
Expands reasoning when deeper diagnosis is needed
These shifts are developer-reported observations and should be verified through reproducible evaluation and studies with real learners and educators.
Responsible release
Limitations, risks & safeguards
Incorrect or fabricated instruction
The model can produce plausible but wrong explanations, examples, or feedback. Learners and educators need verification pathways.
Pedagogical overconfidence
A high benchmark score does not establish that the model recognizes every misconception or knows when to withhold an answer.
Bias and unequal performance
Performance may vary across language, culture, disability, age, subject, and educational context; the draft does not provide disaggregated results.
Privacy and child safety
Local deployment can reduce third-party transmission but does not itself establish appropriate access controls, retention, consent, or safeguarding.
Academic integrity
A tutoring model can still be used to complete work rather than support learning. Deployment should include clear norms and educator controls.
Automation bias
Institutions should not treat generated tutoring judgments as authoritative or use the model for consequential student decisions.
Availability
Run and evaluate it locally
The release draft describes open weights, quantized GGUF variants, a training-recipe outline, evaluation harness, example trajectories, and a lightweight reference deployment.
Ollama
Local model runner
llama.cpp
CPU and GPU inference
vLLM
High-throughput serving
Reference stack
Learner state + local RAG stub
Attribution
Citation, authors & release resources
Suggested citation
Rivera, A.; Chen, L.; et al. (2026). Teach-1.0: Frontier Pedagogical Intelligence, Democratized. AFAI Open Models.
Authors: A. Rivera*, L. Chen*, M. Okonkwo, S. Patel, J. Alvarez, K. Nakamura, T. Brooks, R. Singh
* Equal contribution