~/niansia / notesEN繁简

Research notes

Taiwan Exam design notes: the model writes, the scripts only check

An Agent Skill that has an AI write original GSAT practice exams — why its scripts deliberately never vouch for question quality, and how the checks, locks and template stamping are designed.

3 min readAgent SkillLLMsystem design

Taiwan Exam has an AI write an original practice exam for Taiwan's GSAT and delivers two PDFs: the questions and the worked solutions. Its core division of labour fits in one sentence: the model writes the questions; the scripts only lay out, check and stamp pages — and never vouch for question quality.

Why split it this way

LLMs are good at writing questions, and just as good at saying they wrote them well. If one model writes, self-reviews and declares itself done, quality problems disappear behind the word "checked". So the design rule is:

The main workflow script says so plainly: it "never authors questions or approves reviews".

The pipeline

  1. Load knowledge, preflight. Read only what this subject needs, in chunks of at most 12,000 characters; check the PyMuPDF version, the hashes of the 30 original template PDFs, the answer-rate calibration snapshot, fonts, and whether images can be downloaded.
  2. Paper plan. Lay out the whole paper first: item numbers, points, four difficulty bands, answer-key distribution, how many figures, total time. Difficulty is calibrated to the official answer rates from the 2022–2026 exams; for maths that means "medium-hard + hard ≥ 70 points, hard ≥ 30 points, 80–92 minutes by hand".
  3. Write in batches. Two to four items at a time, each batch checkpointed. Bad formats are refused outright: LaTeX, $ signs, unbalanced tags, figures whose hash does not match, low-resolution images.
  4. Blind re-solving. An answer-free packet is built, each item is solved again from the printed page, and its answer rate is estimated and placed in a difficulty band.
  5. Content lock. Questions and answers are written into a lock file before any real layout starts; changing a question means re-locking with a written reason.
  6. Stamp onto the original templates. The question body is a transparent page, placed onto the original CEEC template with PyMuPDF; only the year, page numbers and cover title are dynamic. Header and footer pixels must be identical before and after stamping, or the job aborts.
  7. Check every page, then deliver. Every page and every item crop needs written observations and a pass/fail mark; only after the final check passes are the two PDFs written.

What the checks look like

Making it run in everyday AI apps

Teachers and students rarely use a terminal, so the same rules ship three ways: a Skill ZIP to upload to Claude or ChatGPT (Claude's uploader accepts at most 200 files, so resources are bundled), a knowledge file for Gemini and project-based setups, and a local install for Codex and Claude Code. A full paper often takes two or three rounds, so every step checkpoints, and replying "continue" picks up where it stopped.

Honest limits

Source and installation: github.com/niansia/taiwan-exam. There is also an 80-second trailer.

More notes

LumiGrid notes: what a luminance-guided curve grid actually learnedFrom a 16.48 dB course pipeline to 24.57 dB — the design, the ablation, two failure cases, and one finding that does not flatter the method.Porting LumiGrid to the browser: a 6-channel convolution that only broke on WebGPUTwo neural networks and a hand-written WebGL shader, verified pixel by pixel against PyTorch — and how a bug that left the WebGPU version at 28.75 dB was found by bisection.
Research notesBack to the portfolio