AFAI logo
Open Models
Research releaseOpen weightsEducation

Teach-1.0

Frontier pedagogical intelligence, democratized.

A specialized 4B-class tutoring model trained for misconception diagnosis, scaffolding decisions, sustained multi-turn teaching, and student-progress signals.

4B class

Parameters

51.8

Reported MainScore

Self-hosted

Deployment

25.08.26

Release

Authors: A. Rivera*, L. Chen*, M. Okonkwo, S. Patel, J. Alvarez, K. Nakamura, T. Brooks, R. Singh
* Equal contribution

Documentation status: This page reflects the supplied release draft. License, language coverage, full data provenance, training compute, detailed evaluation methodology, and independent validation should be completed before production use.

Release summary

Specialized for the work of teaching

Teach-1.0 is optimized for diagnosing misconceptions, deciding when to scaffold versus when to hold back, sustaining multi-turn sessions, and producing measurable student-progress signals.

It is reported to be strongest on K–12 and early higher-education mathematics and STEM tutoring. Quantized deployment is intended to let schools, universities, and lower-resource institutions self-host without continuous API costs or student conversations leaving institutional infrastructure.

The model is derived from a strong open 4B-class base in the Qwen3-4B family and receives additional supervised and reinforcement-learning stages focused on pedagogical behavior.

Model card

Model details

Known release metadata is stated directly; fields not established by the supplied draft remain labeled as such.

Developer

AFAI Open Models

Release

Teach-1.0

Release date

August 25, 2026

Model type

Specialized pedagogical language model

Model size

4B class

Base model

Qwen3-4B family

Modality

Text

Primary domain

K–12 and early higher-education mathematics and STEM

License

Not specified in current release draft

Languages

Not specified in current release draft

Context length

Not disclosed

Release status

Research release / documentation in progress

Evaluation

Pedagogical performance

MainScore is described as a weighted aggregate of process quality, scaffolding decisions, misconception diagnosis, and simulated student-outcome proxies. Constraint-violating solutions receive zero.

ModelScoreNotes
Base Qwen3-4B28.4Strong general reasoner
Phi-4-mini31.7Excellent small reasoner
Larger open 8–14B general39–44Typical prompted baselines
Strong proprietary mid-tier48–52API models under tutoring prompts
Teach-1.051.8Specialized 4B-class model
Evaluation caveat: MainScore is a custom reported metric. The draft does not provide the full harness, sample sizes, uncertainty, contamination analysis, disaggregated results, or independent replication. Simulated outcome proxies are not evidence of improved real-world student learning.

Use guidance

Intended and out-of-scope uses

Intended uses

  • Tutoring assistance with educator or institutional oversight
  • Mathematics and STEM learning support
  • Research on pedagogical dialogue and scaffolding
  • Locally hosted educational prototypes
  • Evaluation and improvement of tutoring workflows

Out of scope without further validation

  • Replacing qualified teachers or human support
  • High-stakes grading, admissions, or disciplinary decisions
  • Safety-critical, medical, legal, or crisis guidance
  • Unsupervised deployment to children
  • Subjects and languages not adequately evaluated
  • Assuming generated explanations are accurate

Development

How Teach-1.0 was trained

01

Stable long-horizon RL

Training techniques preserve policy entropy across longer rollouts and reduce train–inference mismatch, with the goal of avoiding repetitive or overly verbose behavior.

02

Pedagogical data pipeline

Trajectories receive automated and sampled human checks for verifier quality, difficulty calibration, and answer-dumping or reward exploitation.

03

Self-compaction

Near context limits, the model produces a concise working-state summary of student understanding, open goals, and recent misconceptions before continuing.

04

Efficiency pressure

Alternating phases optimize pedagogical success and then apply soft pressure on tokens, turns, and unnecessary elaboration.

Data

Training data & verification

The supplied draft reports a mixture of real de-identified tutoring transcripts and filtered synthetic trajectories generated under pedagogical constraints.

Reported quality controls

  • Verifier false-positive and false-negative checks
  • Difficulty calibration against the base model
  • Answer-dumping and exploit detection
  • Network isolation and grading-path separation
  • Sampled human review

Documentation still needed

  • Dataset names, sizes, and mixture proportions
  • Transcript consent and de-identification process
  • Collection geography and learner demographics
  • Copyright, licensing, and retention details
  • Bias and representation analysis

Observed behavior

Reported behavioral shifts

Explores student reasoning before intervening

Uses more condensed reasoning on routine steps

Probes assumptions and possible misconceptions

Asks targeted questions when student intent is ambiguous

Maintains more consistent help-versus-hold-back decisions

Expands reasoning when deeper diagnosis is needed

These shifts are developer-reported observations and should be verified through reproducible evaluation and studies with real learners and educators.

Responsible release

Limitations, risks & safeguards

Incorrect or fabricated instruction

The model can produce plausible but wrong explanations, examples, or feedback. Learners and educators need verification pathways.

Pedagogical overconfidence

A high benchmark score does not establish that the model recognizes every misconception or knows when to withhold an answer.

Bias and unequal performance

Performance may vary across language, culture, disability, age, subject, and educational context; the draft does not provide disaggregated results.

Privacy and child safety

Local deployment can reduce third-party transmission but does not itself establish appropriate access controls, retention, consent, or safeguarding.

Academic integrity

A tutoring model can still be used to complete work rather than support learning. Deployment should include clear norms and educator controls.

Automation bias

Institutions should not treat generated tutoring judgments as authoritative or use the model for consequential student decisions.

Availability

Run and evaluate it locally

The release draft describes open weights, quantized GGUF variants, a training-recipe outline, evaluation harness, example trajectories, and a lightweight reference deployment.

Ollama

Local model runner

llama.cpp

CPU and GPU inference

vLLM

High-throughput serving

Reference stack

Learner state + local RAG stub

Before deployment: confirm the final license, hardware requirements, model hashes, configuration, privacy controls, logging, age-appropriate safeguards, evaluation scope, and escalation to qualified human support.

Attribution

Citation, authors & release resources

Suggested citation

Rivera, A.; Chen, L.; et al. (2026). Teach-1.0: Frontier Pedagogical Intelligence, Democratized. AFAI Open Models.

Authors: A. Rivera*, L. Chen*, M. Okonkwo, S. Patel, J. Alvarez, K. Nakamura, T. Brooks, R. Singh

* Equal contribution

Research release

Try it. Host it. Stress-test it.

Tell us where the long sessions, scaffolding decisions, and local deployment still fall short.