VCE exam analysis · a kit for teachers

Build your own exam analysis tool

At the end of this you'll have a tool that maps every past exam question against your study design, shows where the marks actually sit, ranks what's most likely to come up next, and builds practice papers to a blueprint. You gather the files and check the coding. Claude does the reading and the building.

≈2 hours, across a few sittings 3 phases, 9 steps Most of it is downloading and verifying
The diagnostic comes first — one minute, three questions. Answer it to unlock the build.
Before you download anything

Does this kit fit your study?

This approach works by finding patterns across a large number of individually-marked questions. Some VCE studies give you that; some don't. Open a recent exam, its examiner report and the study design, and answer these three.

Question 1 of 3

Open a recent exam. How many separately-marked questions are there in the whole paper?

Question 2 of 3

Open the matching examiner report. Does it give you a number for each question — an average mark, or a distribution of how many students scored 0, 1, 2 and so on?

Question 3 of 3

Open the study design. Could you list what's examinable as discrete content points a single question could target?

Verdict

Answer all three to see where you stand

The result isn't softened to be encouraging. A tool built on data that can't support it looks authoritative and says almost nothing, which is worse than no tool.

Facilitator Do this live, together, before anyone opens a browser tab. It takes a minute and it's the difference between an afternoon well spent and a tool that misleads a faculty. Expect English, Literature, EAL and Languages to land on stop — say plainly that a real tool for those studies measures different things (prompt archetypes, criteria against report commentary, stimulus patterns, text list movement) and needs its own worked exemplar. Don't offer to adapt this one.
Phase 1

The project

A Claude Project holding your study's official documents. Everything downstream is read out of these files, so this is the only phase where nothing can be automated — it needs your judgement about what's current.

Step 1

Gather the source files

On the VCAA site: Study Designs → your study, and the Past examinations page.

≈30 min
FileWhy it matters
Study design (current)Becomes the key knowledge index — the spine of the whole tool
Support materials / FAQsClarifies grey areas. Often answers exactly what students get wrong
Specifications, conventions, glossariesSubject-specific rules a question writer must respect
Sample exam + answersThe intended shape of the paper before any live exam existed
Past exam papers — every year under the current study designThe question corpus. Every one gets coded
Insert / resource / stimulus booksEssential if your exam supplies stimulus. Without them the papers don't parse
Examiner reports — every matching yearWhere the performance data comes from. Skip one and that paper's questions have no average
Current study design only Older papers were written against a syllabus that no longer applies, and mixing them in corrupts the coverage analysis — you'll get key knowledge points that look well-assessed but were assessed under different wording. One or two compliant papers is still workable: the tool leans harder on the sample exam and reports honestly that the evidence base is thin. Say so rather than padding it out with superseded papers.
The failure that hides itself Scanned PDFs won't work, and they fail silently — Claude simply doesn't see those questions, and never says so. Open each file and try to select text with your cursor. If you can't select it, it's a scan: run it through OCR (Acrobat, or free tools like OCRmyPDF) and upload that version instead.

Don't take the selectable-text test on trust. Before you upload thirty files, upload one exam paper to a chat and run this:

Pre-flight check · paste into any chat with one paper attached
I've attached one VCE {{STUDY}} exam paper. Don't analyse it yet.

Tell me only this:
1. Can you read the text of this PDF, or is it a scanned image?
2. Quote the exact wording of the last separately-marked question in the paper, and its mark allocation.
3. How many separately-marked questions are in the whole paper, and what do the section totals add to?

If you can't read any part of it, say which pages.

If it can't quote the last question back to you accurately, the file is unusable and so is every other file that looks like it. Fix them all before Step 2.

Name files so the year and type are obvious2025-Biology-exam.pdf, 2025-Biology-report.pdf. Claude cites by filename, so this is what makes its coding checkable later.

Facilitator This is the step people do at home, and the step where the whole build quietly dies. If you have time in the session, get everyone to run the pre-flight check on one paper while you're in the room — a scan discovered now costs five minutes, a scan discovered in Phase 2 costs the afternoon.
Step 2

Create the project and upload

Your own project, for your subject only.

≈10 min
  1. In the Claude sidebar, open Projects, then Create project.
  2. Name it for your study — VCE [your study] Exam Analysis.
  3. Upload everything from Step 1 into the project's knowledge.
The reference tool comes later You don't need it yet. It's a finished tool for a different VCE study, which Claude reads as the structural model when you build yours in Phase 3 — a working exemplar beats any amount of description in a prompt, so it's the single most reliable thing you can do. It's in the handover pack you'll download at the end of Phase 2, along with your filled-in project instructions.
One subject per project Don't build in the shared faculty project, and don't share yours with a colleague teaching a different study. Projects don't partition: if Biology papers and Legal Studies papers sit in the same project, Claude reads both at once with no way to tell which study it's working on — and it will happily code a question from the wrong subject.
Facilitator Screen-share this one rather than describing it. The two things people get stuck on are finding Projects in the sidebar, and the difference between attaching a file to a chat message and adding it to the project's knowledge. Show both, and say which one this is.
Step 3

Paste in the project instructions

In your new project, open its instructions and paste this in whole. It's what holds the accuracy line for every message that follows.

≈2 min
Project instructions · your project
ROLE
You are an experienced VCE {{STUDY}} teacher, exam assessor and data analyst,
working with a teacher who is building an exam analysis tool for their faculty.

YOUR SOURCES
Everything you produce must be grounded in the files in this project:
- The current study design is the authority on what is examinable and on
  correct terminology. Never introduce content outside it.
- The examiner reports are the authority on cohort performance and on what
  loses marks. Draw on them heavily.
- Past exam papers are the corpus. Quote question stems accurately.
Cite the file and year for every specific claim. If the files don't tell you
something, write UNKNOWN rather than guessing.

ACCURACY RULES — these matter more than fluency
- Never invent a statistic, an average mark, or a question that doesn't exist.
- Never estimate a figure that could be read from a file. Read it.
- When you tag a question against key knowledge, you must be able to point to
  the words in the question stem that justify the tag.
- Flag your own uncertainty inline rather than presenting a guess cleanly.

WHAT NOT TO DO
- Don't reproduce an entire past exam paper on request.
- Don't state that a future exam will contain something. Rotation patterns are
  hypotheses about likelihood, never forecasts, and the whole study design
  remains examinable every year.
That last rule matters. Your own examiner reports warn students against relying on past papers — the tool has to hold that line, or it teaches exactly the behaviour VCAA is cautioning against.
Phase 2

The dataset

This is the real work, and it's where your subject expertise earns its keep. Three passes, each verified before the next. The dataset is the product — the interface in Phase 3 is only a window onto it.

Step 4

Build the key knowledge index

The spine of everything else. A gap here becomes an invisible blind spot.

≈15 min
Prompt · Step 4
Read the study design and produce a complete index of examinable key knowledge
for Units 3 and 4, as a JavaScript object.

Format each entry as:
  "CODE": { o:"U3 AoS1", label:"short plain-language description" }

Use a code scheme of [Unit][Outcome]-[number], e.g. "U3O1-1". Add separate
codes for any cross-cutting terminology the study design defines separately
and that questions could target directly.

Rules:
- Cover every key knowledge point in Units 3 and 4. Completeness matters more
  than brevity — this index is the spine of everything else.
- Keep labels under about 15 words, in language a student would recognise.
- Cite the study design page number for each entry.
Save this index to a file — a plain .txt is fine. You'll load it into the Tag Checker at Step 5, which is what turns a code like U3O2-4 back into words you can actually check a question against. Save each paper's coded questions to its own file as you go, too.
Tick both checks first.
Step 5

Code every question

One paper at a time — doing all of them in one message produces shallower tagging. Then check the tags yourself. This is the highest-value hour in the whole process.

≈25 min per paper
Prompt · Step 5 · run once per paper
Code every question in the 2025 {{STUDY}} exam as JavaScript objects in this
shape:

{ paper:"E25", sec:"A", q:"1", marks:4, cmd:["Identify","Discuss"],
  kk:["CODE1","CODE2"], avg:2.3, dist:[19,9,17,35,20],
  stem:"the question stem, quoted accurately",
  why:{ "CODE1":"the words in the stem that justify this tag",
        "CODE2":"the words in the stem that justify this tag" },
  choice:"what the student gets to choose, if anything" }

- cmd: every command term in the question, in the order they appear.
- kk: the key knowledge codes from the index that the question actually
  assesses. Usually one to three. Only tag what the question truly requires,
  not everything it touches — over-tagging flattens the coverage analysis.
- avg: the average mark from the examiner report for that question. Read it
  from the mark distribution table. If the report gives a distribution rather
  than an average, calculate it and show your working. Omit avg entirely if
  the report doesn't cover this paper.
- dist: the mark distribution from that same table, as an array running from 0
  marks up to full marks — so a 4-mark question has 5 numbers. Use the report's
  own figures, usually percentages of the cohort. Omit dist if the report gives
  only an average. You are already reading this table for avg, so capturing it
  costs nothing and it is what makes avg checkable later.
- why: one entry per kk code, quoting the words in the stem that justify that
  tag. Quote the stem, don't paraphrase it. Every code in kk needs an entry.

Use exactly these field names. Don't add fields I haven't asked for.

Change the year each time you run it. Keep track here — a paper isn't done when it's coded, it's done when it's verified.

Don't skip this next part Checking tags is slower and duller than watching a tool appear, and this is where nearly everyone stalls. Every error you catch here is an error that would otherwise have shipped invisibly — a beautiful interface over a badly coded dataset produces confident, wrong answers about where the marks sit, and nobody notices, because it looks authoritative.
Do this in the Tag Checker The Tag Checker is the other half of this kit. Drop in the index and every question file — it works out which is which — and the dataset becomes a proof sheet: codes resolved to their labels, justifications beside the stems, and every mechanical error found for you. It flags; it never fixes.

You can change the tags in place there: codes are picked from your own index, so a code that doesn't exist can't be created, and every tag change needs a reason — which then becomes the stored justification. There's a tick per question, so you can stop at 23 and come back. It also marks which questions can't be judged without the exam paper in front of you, and filters to just those. When you're done it writes a corrections file to paste back into your project.

Nobody needs a Claude account to use it. That makes this step delegable: send a colleague the dataset files, the papers and that link, and they send back the corrections file. Two people checking a paper each is closer to how you'd moderate anything else — and it beats the person who did the coding being the only one who ever reads it.

Once the flags are clear, the part no checker can do:

Facilitator If you do one thing in the live session, do this: code one paper together on the screen, and find a wrong tag in front of the room. People need to see that the tagging is genuinely arguable before they'll believe it's worth checking. Say out loud that this hour is the difference between analysis and decoration.

A good sequence: open tagcheck.jotty.au, drop in a dataset with a planted fault, let the room watch every mechanical error appear at once — then close the laptop lid on that and find a wrong tag by eye. The contrast is the lesson. The machine half is instant; the judgement half isn't, and that's the hour you're asking them for.
Tick all five checks first. This is the gate that matters.
Step 6

Command term glossary and report findings

A generic definition of Explain is available anywhere. What a full-mark Explain looks like in your study, according to your assessors, is not.

≈10 min
Prompt · Step 6
1. From the coded questions, produce a command term glossary:
   Term: { marks:[min,max], def:"generic definition",
           note:"what the examiner reports specifically say about how this
                 term is handled in THIS study, with the year cited" }

2. From the examiner reports, list every weakness they name across all years,
   with the year and the exact question it relates to. Mark which ones recur
   in more than one report — those are the persistent ones.

3. List the terminology confusions the reports name directly. Search for
   "confused", "incorrectly", "misunderstood", "uncertainty around".

Read the note field on a few terms. If it's a generic definition restated, push back and ask for the report wording and the year — that field is the reason the glossary is worth having.

Phase 3

The tool

A single self-contained HTML file your faculty can open, share and print from. It's only as good as the dataset underneath — which is why it comes last.

Step 7

Build it in stages

Don't ask for the whole thing at once. Shell plus two views, check them, then add views one per message.

≈20 min
Before you run this Upload reference-exam-analysis-tool.html — from the handover pack you downloaded at the end of Phase 2 — into your project's knowledge. The prompt below tells Claude to read it as the structural model, so without it there this step produces something generic.
Prompt · Step 7 · first build
Using reference-exam-analysis-tool.html as the structural model,
build the equivalent tool for VCE {{STUDY}} as a single self-contained
HTML file — no external dependencies, no browser storage.

Use the verified dataset we built: the key knowledge index, the coded
questions, the command glossary and the report findings. Do not add data
that isn't in it, and do not smooth over gaps.

Start with the shell plus two views only:
  - Overview: headline findings, papers included, what's been assessed in
    every live paper and what's never been assessed at all.
  - Question explorer: every coded question, filterable by paper, section,
    command term and key knowledge.

Match the reference tool's approach to layout and interaction, but derive the
visual treatment from {{STUDY}} rather than copying the reference's palette.

Then add the remaining views one per message. Each has its own prompt below — open it, copy it, paste it, check what comes back, then move to the next. Which ones you should build depends on your diagnostic result, and a view marked Restrict has that limit written into its prompt already.

This is a teacher's tool Don't build a student view of it, and don't hand the file to a class. Students given a ranked likelihood list will study to it — they'll revise the top five and drop the rest, which is the opposite of what the analysis is for. Use it to decide what you teach and in what order; what reaches students is your teaching, and revision advice you've already filtered through your own judgement.
Step 8

Test it against things you already know

Four checks. You are the only instrument that can run them.

≈15 min
The rule for anything you can't trace If the tool states a figure you can't trace, ask it which file and page it came from. Anything untraceable comes out — not softened, not caveated. Out.
Run all four checks first.
Step 9

Share it, and keep it honest

It's a single HTML file. Email it, drop it in a shared drive, or share the project so colleagues can regenerate it.

≈5 min

Each year, when VCAA publishes the new exam and report: add both to the project, code the new paper with the Step 5 prompt, and rebuild. Then compare the new coverage matrix against last year's — what rotated in and out is the most useful thing the tool will tell you.

Next year, and the year after The annual re-verification is the step most likely to get skipped, because by then the build feels finished. It's also the easiest to share out: send the new paper's dataset and the exam to a colleague with the Tag Checker and they can do the whole check without a Claude account, in the time it takes to mark a set. Their corrections file goes into your project and the tool rebuilds. Half an hour a year is what keeps this honest.
Two things to say out loud when you hand it round The predictions are pattern-reading, not forecasting. The whole study design is examinable every year and every paper is unique. The tool is for prioritising revision when time is short, not for narrowing what gets taught.

The analysis is only as good as the coding. If a colleague disagrees with a tag, they're probably right to check — send them back to the dataset rather than defending the interface.

What generalises, and what doesn't

Scenario generation

The reference uses fictional client briefs with design criteria and communication needs. Your equivalent might be case scenarios (Legal Studies), stimulus datasets (Biology), or a business situation (Business Management). Same idea, different shape.

Student choice

Some studies let students choose which context, case or option they answer on. If your exam has a choice like that, code it — it changes what the coverage matrix means. If not, drop the field entirely rather than leaving it empty.

Practical components

If your study has a SAT or performance component, the exam only covers part of the story. Say so in the overview rather than letting the tool imply it's the whole picture.

The temptation

To jump to Phase 3. The dataset is the product; the interface is only a window onto it. Nothing in Phase 3 can rescue a Phase 2 you rushed.

Your study name, progress and diagnostic answers are stored in this browser only — nobody else sees them, and clearing your browser data clears them.

The other half of this kit is the Tag Checker, where the Phase 2 verification happens.