LittleStory
A voice-first documentation app that turns the teacher's day into tomorrow's plan.
Problem
At the end of the day a teacher still has paperwork left: the observations they carried in their head all day have to be written up at night, at home, by hand. Documentation and administration take a large share of educators' working time, and heavy paperwork is a documented reason early-childhood staff consider leaving — while the tools that promise help often add photo quotas, tags and points, turning the classroom into data work. Meanwhile collecting data about young children raises rights concerns, so the safe place for AI is beside the teacher, never in front of the child.
Constraints
- The child never interacts with the app: no child account, no child-facing screen, no data collected from children. The teacher's voice is the only input.
- Nothing is saved to a child's record without the teacher reviewing and confirming it — the trust gate is the product, not a feature.
- The system must not diagnose, rank or determine what a child has mastered; the model drafts, the teacher decides.
- Bilingual English and Bahasa Indonesia end to end — interface, speech input and AI output — for the intended pilot with PAUD/TK teachers.
- Classrooms are offline and audio is expensive, so the app keeps a local snapshot with an outbox and retries.
Decisions
-
One button, no form: the teacher speaks freely about the day and the app does the structuring.
rejected A structured observation form with fields, tags and per-child checklists.
Forms move the work from the end of the day to the middle of it and ask the teacher to think in the system's categories while thirty children are in the room. Speaking about the day the way you would tell a colleague needs no training and no screen, which is the only input method that survives a real classroom.
-
Put a review gate between the AI draft and the child's record, with every field editable.
rejected Auto-commit the structured output to the child's record once the model is confident.
A model that is wrong about a child is not a bug report, it is a record that follows them. Making the teacher the author of the final text keeps professional judgement in charge and gives the model a role it can be honest about: a draft that reduces typing, not an authority that assigns meaning.
-
Drive language with one preference that follows the whole pipeline — interface, speech decoding and AI output together.
rejected Translate the interface and leave the speech and model in English.
A bilingual UI over an English-only pipeline is worse than a monolingual app, because the teacher reads Indonesian and gets structured output in a language they did not speak. Whisper's decoding language and the structuring prompt's output-language directive have to travel with the same preference, and the stored values stay English data so the records remain machine-readable while only the display is translated.
-
Build the safety layer as code, not as prompt text — shielding diagnostics and ranking, bilingual safeguarding signals, and a hard rule about what the model may not produce.
rejected Instruct the model in the system prompt not to diagnose children.
A prompt instruction fails silently and cannot be tested. Keyword and category shields in both languages, machine error codes on failed jobs and quality gates on the generated structure are things a test suite can assert; the safeguarding and diagnostic keyword lists are the union of the English and Indonesian terms precisely because a child's risk does not arrive in only one language.
The hard part
Building a language model into an early-childhood workflow without letting it sound like an assessment. Everything the model produces is a draft over a teacher's words, and the failure mode is not a spectacular error — it is plausible, fluent, slightly wrong text about a child that a tired teacher accepts because it arrived pre-written. The work that mattered was refusing affordances: no mastery score, no developmental ranking, no red flags a model could raise on its own, and a visible way back to the teacher's own words for anything it produced.
Outcome
-
178 / 78
app and server tests green in the project's last recorded cycle
-
2
languages end to end — interface, speech input and AI output
-
0
child-facing screens or child accounts
-
Working prototype
teacher study planned as future work, per the submission
Figures and revisions
source: LittleStory local prototypechecked: 11 Oct 2026
| plate | source | checked |
|---|---|---|
| FIG. 1 | LittleStory local prototype | 11 Oct 2026 |