AI tutors grounded in your curriculum
What should you know about AI tutor development?
A grounded tutor answers only from approved content, cites the lesson and hands off to a teacher when out of scope. AI tutor development is mostly retrieval, pedagogy rules and evaluation, not model choice: the tutor must explain rather than answer, stay inside the syllabus and prove on a test set that it does both.
AI tutor development goes wrong in a predictable way. A team connects a chat model to a lesson library, the demo is charming, and then a student asks about a topic two chapters ahead, or a topic not on the syllabus at all, and the tutor answers fluently and incorrectly. Parents notice. Teachers stop trusting it. The fix is not a smarter model; it is grounding, pedagogy rules and evaluation, designed in from the first week.
This article explains what a grounded edtech AI tutor is, how it differs from a general chatbot, what the architecture looks like, how to make it teach rather than tell, and what it costs to build. It is written for edtech founders, school-network CTOs and university learning-technology teams considering an LMS AI assistant for their own content.
What a curriculum-based AI tutor is
A curriculum-based AI tutor is an assistant whose every answer is derived from content your institution has approved: textbooks, lesson plans, worked examples, past papers and teacher notes. It knows which unit the student is in, answers with a citation to the lesson, adapts the explanation to the student's level, and refuses politely when a question is outside the approved material, offering to notify a teacher instead. That last behaviour is what distinguishes it from a consumer chatbot and what makes it acceptable to a school.
General chatbot vs grounded tutor
| Property | General chatbot on a model API | Grounded curriculum tutor |
|---|---|---|
| Source of answers | Model's training data | Your approved content, retrieved per question |
| Citations | None or invented | Lesson, page and section for every explanation |
| Scope control | Answers anything | Declines out-of-syllabus questions and hands off |
| Pedagogy | Gives the answer | Guides with hints and checks understanding before revealing |
| Level awareness | None | Uses the student's grade, unit and recent progress |
| Teacher visibility | None | Every conversation and hand-off visible in the LMS |
| Evaluation | Anecdotal | Golden set of student questions with expected behaviour, run on every change |
Grounding: the tutor answers only from approved content
Grounding is the same retrieval discipline we use for enterprise knowledge systems, applied to lessons. Content is parsed with structure preserved (a worked example must stay with its problem, a diagram caption with its diagram), chunked by lesson section rather than by fixed token windows, and tagged with subject, grade, board, unit and language. At question time the tutor retrieves only from the units the student has been assigned, plus any prerequisites, and composes an explanation from what it retrieved, citing the section. If nothing relevant is retrieved, it says so rather than improvising.
Two details matter more than they look. First, scope is a filter applied before ranking, not an instruction in the prompt, because instructions can be argued with and filters cannot. Second, the same question can be in scope for one student and out of scope for another, so retrieval is per student, not per subject. The engineering is described on our retrieval and knowledge engineering page, and the failure modes in why basic RAG fails in production.
Pedagogy rules: teach, don't tell
A tutor that hands over the final answer is a homework machine, and schools will not buy one. The tutor needs a small set of pedagogy rules that govern how it responds, enforced by the application rather than left to the model's mood. Typical rules we implement:
- For a practice problem, give a hint first, then a worked step, then the answer only after the student attempts it
- Ask one check-for-understanding question after each explanation, and adapt the next explanation to the answer
- Use the vocabulary and notation from the student's textbook, not the model's default
- Keep explanations to the student's grade level, with a simpler variant available on request
- Never mark an assessed assignment; direct the student to the teacher for graded work
- Switch to the student's preferred language while keeping technical terms consistent with the syllabus
These rules are tested like any other behaviour: a set of student questions with the expected shape of response (hint, not answer) is part of the evaluation suite, and a change that starts revealing answers fails the build.
Hand-off to a teacher
Out-of-scope questions, signs of frustration, repeated wrong answers on the same concept, and anything touching wellbeing go to a person. The hand-off carries context: the student, the unit, the last few turns and the tutor's own assessment of the difficulty. Teachers see this in the LMS or in a daily digest, so a hand-off becomes a teaching signal rather than a support ticket. The pattern is the same one we use for customer service agents: escalate with context, never with a blank slate.
Integration with the LMS
An LMS AI assistant needs to know who the student is, what they are assigned and what they have done. Most platforms expose this through LTI (Learning Tools Interoperability) and the associated services for roster and assignment data, which is why we build the tutor as an LTI tool rather than a separate app; the specifications are published by 1EdTech via the IMS Global GitHub organisation. For platforms without LTI, a small integration on top of the platform's API does the same job. The tutor writes back a summary of each session and any hand-offs, so the teacher's view of the student stays complete.
Evaluation: proving the tutor behaves
Before any student sees the tutor, we build a golden set of a few hundred real or realistic questions across subjects, grades and languages, each with the expected behaviour: the section it should cite, whether it should hint or explain, whether it should decline. We measure retrieval recall (did the right section appear), groundedness (did the explanation stay inside the retrieved text), pedagogy compliance (hint before answer where required) and scope compliance (declined when it should). Teachers review a sample of transcripts every week during the pilot. Our post on evals as the practice that separates demos from products explains the discipline in general.
A worked example
A test-preparation edtech company with its own video lessons and question banks wanted a tutor inside its app. The first attempt, built in-house on a model API, was popular and dangerous: it answered questions from outside the syllabus, occasionally with errors, and gave away answers to practice sets. We rebuilt it as a grounded tutor: lessons and question explanations parsed and tagged by unit, retrieval scoped to the student's current and completed units, pedagogy rules that enforced hint-first responses on practice problems, and a hand-off to the company's doubt-solving team with the transcript attached. The evaluation set was built with the company's subject experts and run on every content or prompt change. Teachers reported that hand-offs arrived with enough context to answer in a sentence, and the content team used the "declined" log to find gaps in the lesson library. The approach mirrors the in-app copilot pattern we use in SaaS, with pedagogy rules in place of business policy.
Team and timeline
A grounded tutor for one subject and grade band is typically an AI engineer for retrieval and evaluation, a full-stack engineer for the LTI integration and interface, and a subject-matter reviewer from your side, over eight to twelve weeks. The first three weeks parse content and build the golden set; the middle weeks tune retrieval and pedagogy rules against it; the last weeks integrate with the LMS, run a teacher-supervised pilot and hand over. Where the content is unstructured or spans many boards and languages, a three-week ProofRun on one subject first is the cheaper way to find out. LLM application builds start from $21,000 / ₹13.6L and retrieval work from $14,000 / ₹8.8L; current figures are on the pricing page, and our education sector page lists related work.
Before you start: a checklist
- Inventory your content by subject, grade, board and language, and note what is missing
- Decide the pedagogy rules with teachers, in writing, before any prompt is written
- Define scope: which units a student may ask about, and what happens outside them
- Confirm how the LMS exposes roster, assignment and progress data
- Collect a few hundred real student questions for the golden set
- Agree who receives hand-offs and how quickly they respond
- Plan the consent and data-retention position for minors before the pilot
Questions clients ask
- Can the tutor use a general model at all? Yes, for language and explanation; it just may not use the model's memory as a source. Every fact comes from retrieved content.
- Which model should we use? We route per task across OpenAI, Anthropic, Google and open-weight models and benchmark on your golden set; the model matters less than grounding and evaluation.
- Does it work in Indian languages? Yes, provided the content exists or is translated and reviewed in that language; the tutor should not translate the syllabus on the fly.
- What about student data? Minors' data needs verifiable consent, strict access scopes and retention limits; see the learner data protection guide below.
Related reading
Read learner data protection: minors, consent and access before the pilot, adaptive learning paths for the sequencing engine a tutor pairs with, and how to measure RAG quality for the retrieval metrics that make the tutor trustworthy.
A tutor earns a place in the classroom by staying inside the syllabus, showing its sources and knowing when to call the teacher.
Frequently asked questions
How is a grounded AI tutor different from ChatGPT?
▾
It answers only from your approved content with citations, follows pedagogy rules such as hint-before-answer, declines out-of-syllabus questions and hands off to a teacher with context. A general chatbot does none of these reliably.
How long does AI tutor development take?
▾
Eight to twelve weeks for one subject and grade band, including LMS integration and a teacher-supervised pilot. A three-week proof of concept on one subject is the sensible first step when content is messy.
Can the tutor be stopped from giving away answers?
▾
Yes. Pedagogy rules are enforced by the application and tested in the evaluation suite, so a change that starts revealing answers to practice problems fails before release.