AI agents assignments for students survive a model doing them when the grade rests on evidence the grader can check independently: by rerunning what was submitted, by comparing it with a key the student never saw, or by asking the student about it for ten minutes. No file a student uploads is proof by itself.
“Survive” has a narrow meaning here. It means the grade rests on something the grader verified and did not take on trust. Only the rows with a live or in-room check also say something about the student; the rerun rows grade the instrument, whoever wrote it. A page that promises assignments a model cannot do is selling a design with a short life.
This post is a bank of ten AI agents assignments for students, each with what is submitted, what the grader does, how it can still be gamed and what it costs in staff minutes, plus a rubric skeleton, a syllabus paragraph on AI tools and an audit list. The order of the course and its assessment plan are in the companion post on the AI agents course syllabus.
One caveat about my evidence. In my searches on 2026-10-06 I found no public thread in which an instructor of an agents course asks how to grade it, so what follows transfers lessons from programming courses and from three published course pages on agents.
Why can no single file prove the work is the student’s?
No file proves it because every kind of file a student can upload, a model can write, including the account of how the work was done. The best public record I know of this is a report from a SIGCSE Virtual working group on a two-hour online workshop held on 28 July 2026, which 73 people joined (64 participants and nine working-group members).
The report is careful about what it is: “It is not a research paper,” and “Nothing here is measured, sampled, or controlled.” Its participants chose to attend, so they are not a sample of instructors (Akbar et al., 2026). Three passages matter for assignment design.
First, logs and reflections. Asked what they use such evidence for, participants in one room “were unanimous that they do not grade it.” The report gives their reasons: “Students omit uses, use tools their peers cannot access, interpret ‘AI use’ inconsistently, or disclose only the low-risk uses.”
The second is the report’s wider conclusion, from a later section: “Nobody in the workshop identified process evidence that resists being generated.” The third is about clever task design. One participant said AI-resistant programming assignments are no longer possible, and the facilitator amended that they are possible “but temporary, lasting perhaps a semester.”
What is the difference between a log and a record?
A log the student writes is testimony about the work, and a recorded run that the grader can replay is a record of it. The distinction comes from the book’s chapter on observability, which defines a trace as the complete, structured record of one agent run and then warns about the part of it the model narrates.
In a trace, the tool calls and results are ground truth about what happened; the narration is testimony about why. Read the testimony; verify against the record. (Chapter 15, “Traces and Spans”)
The book says this about a model’s account of its own reasoning; applying it to a student’s account of their own process is my extension. A prompt log, a reflection and a video are testimony, and a test suite or a recorded run can be executed again.
What can a grader check independently?
A grader can do three things that do not depend on trusting the submission: rerun it, read it against a key, or talk with its author. Those three verbs are the last column of the bank below, because the action decides the cost.
Rerun means a script executes the submission against material the student has not seen. Read means a person compares a short written answer with a key or a second reader’s verdict. Talk means a person spends minutes with the student, live.
Each catches less than its name suggests. A rerun catches numbers that are far from what the system does, and it cannot tell who wrote the code. Reading against a key catches a diagnosis that sounds right and is wrong. Only talk can catch borrowed understanding, and a prepared student can pass it.
Which AI agents assignments for students can a grader check? The assignment bank
These ten AI agents assignments for students each give the grader something to check: five start from an assignment a named course or paper describes, and five are untried proposals. The Origin column separates what each source’s page states, as read on 2026-10-06, from what I added. The staff minutes are my planning estimates, and the gaming column is what I would try as a student with a capable model and a weekend.
| Assignment | Origin | Student submits | What is graded, and what stays hidden | How it can still be gamed | Staff minutes per student | Grader’s action |
|---|---|---|---|---|---|---|
| 1. Break your own loop | Proposal. Michigan EECS 498’s build rubric states the rule that tests earn credit for detecting plausible defects; the same rubric says “No hidden suite,” so hiding the defects is my departure from it | A minimal agent on a scripted model client and the student’s own tests, written to a harness interface the handout fixes | The share of defects the student’s tests catch in the grader’s broken copies. The defects are yours to choose and stay unpublished | A model writes an exhaustive test file from the handout, and a stronger tool writes a better one. The row grades the test file. Use it as low-weight practice | Near 0 after setup | rerun |
| 2. Evaluation suite for a supplied agent | Starts from Stanford CS329Z homework 2, whose page asks for code-based graders, at least one judge, benchmark tasks and error analysis for a pre-built agent. Reference solutions, repeated runs, pass rates and the hidden-variant check are mine; three runs per task is Michigan’s rule | The suite, run records, per-task pass rates over three runs, a pooled rate with its interval, and an error analysis | For each hidden defective variant, at least one task the reference agent passes three times of three and the variant fails three times of three; the pooled rate within the stated tolerance of the rerun; the error analysis read against the failures in the rerun | A model writes and runs the suite, and the numbers then reproduce. Everyone has the same agent, so suites can be shared. An invented rate near the truth passes the tolerance | 5 to read the analysis | rerun, read |
| 3. Teardown of your own agent | Starts from Michigan EECS 498’s teardown report, “graded on the quality of the failures you found.” The recorded runs, the reference solutions, the first-wrong-step criterion and the reproduction check are mine | Three tasks the student’s own agent fails, each with three recorded runs, a reference solution that uses the agent’s own tools, a diagnosis and the command that reproduces it | Whether each failure recurs in at least two of three reruns; whether the diagnosis names the first wrong step in the recorded run. Nothing is hidden | Tasks chosen to fail by construction reproduce every time; the reference solution narrows that route and does not close it. A model writes the diagnoses, and no key exists for them | 8 | rerun, read |
| 4. Three runs, one criterion | Proposal, from the book’s three-run ritual | Three transcripts of one task on which at least one run is rejected, a verdict for each, and the written acceptance criterion | Whether a second reader, given only the criterion and the transcripts, reaches the same three verdicts | Three verdicts match by chance one time in eight. Classmates acting as second readers can agree beforehand. A model drafts the criterion. Use it as a warm-up | 5 to 10; cap the transcript length | read |
| 5. First wrong step, on paper in class | Proposal; related to the error-injection quizzes Engelhardt describes in a physics course | In twenty minutes of class, on paper, for three recorded traces drawn from a pool: the first wrong step, the name of the failure, a fix and one eval case | Agreement with a hidden key on the first wrong step; whether the eval case has a reference solution | Set as a take-home, a model reads the trace, so do not set it as one. The pool leaks between sections and between terms | 6 | read |
| 6. Bug-seeding exchange | Proposal | Round one: a clean harness, a copy with three planted defects and a sealed test for each. Round two: for an anonymized classmate’s planted copy, one failing test per defect found | Planter and finder alike: each test must fail on the planted copy and pass on the clean one. Packs stay anonymous throughout | A model plants trivial defects and writes the sealed tests; planter credit does not depend on subtlety. A model also reads the partner’s copy. Friends can identify each other’s packs. Use it as practice | Near 0; budget time for disputes | rerun |
| 7. Calibrate a judge | Proposal, starting from Stanford CS329T homework 1 (Fall 2025), which has students hand-annotate traces of a provided agent and analyze where judge metrics agree or disagree. The label count, the samples, the statistics and all the checks are mine, following Lab 7 of the kit | 30 hand labels (20 on outputs drawn for that student, 10 on items common to the class), the judge’s verdict on each, the rubric, agreement statistics and an analysis of the disagreements | Statistics recomputed from the table; labels on the ten common items against a hidden staff key; the analysis; a live relabel of five items with one disagreement explained | Copy the judge’s verdicts as hand labels and flip a few, or have a second model label. Nobody checks the twenty individual labels. Five relabels cannot separate an honest student from a prepared one | 5 to read, 5 in person | rerun, read, talk |
| 8. Ten-minute defense | Reported: Stanford CS329Z holds a ten-minute quiz after each homework; Engelhardt describes a one-on-one defense of two or three instructor-chosen regions. The live change is the workshop report’s design principle, which neither source states | Whatever the previous homework produced | Explanation of regions the examiner picks; one change made or planned live. Regions and change stay hidden | Study the generated work, ask a model for the likeliest questions and changes, rehearse. The check catches the student who never read the submission | 10, plus changeover | talk |
| 9. Extension on a clock | Reported format: Michigan EECS 498’s hackathons are new work on the student’s own codebase, in the room, from a prompt revealed there. Its syllabus does not say how they are graded; hidden tests are my proposal | New work on the student’s own codebase, done in the room in a fixed slot, with tools allowed, to an interface the handout fixes | Hidden tests for a requirement revealed in the room | The tool is in the room, so the requirement can be pasted into it: the row shows that a student can operate their own codebase under time, with a tool. A later section learns the requirement | Near 0 per student; one proctored session and one requirement with tests per section | rerun |
| 10. Ship or hold | Proposal | For a supplied change to an agent and supplied before-and-after run records: pooled pass rates with intervals, a comparison of two runs, cost per run and a one-page decision memo | The rates against the grader’s key, computed once; the memo; a ten-minute conversation in which one fact changes | A model drafts the memo and lists the facts that could move; the student rehearses. Whether the decision follows from the evidence has no key | 5 to read, 10 in person, plus changeover | read, talk |
Rows 1, 5, 6, 8 and 10 need no live model from the student. Rows 2, 3, 4 and 7 need one, which can be a model the student serves or one the course provides, and row 9 uses the course’s endpoint in the room. The bank is not in teaching order; sequencing belongs to the design of an agentic AI curriculum as a whole.
What do the rerun rows actually grade?
Rows 1, 2, 3 and 6 grade an artifact by how it behaves when the grader runs it, and they say nothing about who wrote it. The rule behind row 1 is Michigan’s: “Tests earn credit for detecting plausible defects. A correct failing test can retain test credit” (EECS 498 syllabus, Fall 2026). Its rubric for that build also says “No hidden suite”; hiding the defects is my departure from it.
That makes these rows practice with a score, and I would weight them lightly: a stronger tool writes a stronger test file, so a heavy weight partly grades the tool. Their value is exact feedback at almost no cost per student.
Lab 1 of the kit, published with lecture 2, is row 1 graded a new way. Chapter 3 says of the loop’s history that “whatever your code appends is the agent’s memory, and whatever it fails to append never happened.” That handout already asks students to test a dropped append and a model that never says done, so those two are examples to show and not defects to hide.
For row 6 the fit is Lab 2, published with lecture 3. Rows 2 and 7 follow Lab 7, which the kit’s week 8 describes for an eval set on the student’s own agent, and they need a harness the grader can run. The Stanford page behind row 2 is CS329Z Logistics, Fall 2026.
Which rows say something about the student?
Rows 5, 7, 8, 9 and 10 do, because a person watches part of the work or the work happens in the room. The workshop report states the principle: “high-stakes credit should rely primarily on work that can be directly verified, while take-home assignments can function as practice for those assessments.”
Michigan’s syllabus is the clearest example I found in an agents course. Its first hackathon falls “twelve days before your build is due,” on new work “from a prompt revealed in the room,” and within the first phase it counts for half the phase grade, twice the take-home build.
The fifth row applies Chapter 15’s method: “read the trace forward to the first step where something is wrong—not the last step, where the wrongness became visible—because that first step is the bug, and everything after it is consequence.” The agent bug bestiary gives students names for what they find. Reading a run step by step and grading its path is agent trajectory evaluation, and rows 3 and 5 are its classroom form; the kit’s tracing lab is a paragraph in the syllabus, with no lecture page yet.
One caution on row 7’s live relabel. A student who agrees with a model’s labels 83 percent of the time still matches at least four of five with probability 0.80 (5 × 0.83⁴ × 0.17 + 0.83⁵; illustrative). Treat the relabel as the opening of the conversation, not as a scored test.
Preparation passes every talk row, as the report says: “Some students perform well in an oral format by studying AI-generated work carefully, without possessing the underlying skill to have produced it.”
How do you grade a system whose runs vary?
Grade the deterministic part exactly and the model’s part as a rate over several runs, and state in advance how far your rerun may land from the student’s number. The split is the book’s; the tolerance is my proposal. It is the classroom version of how to test AI agents when no two runs match.
The loop, the tools and the parsers are ordinary code. Chapter 15 says to test them against a scripted model, “a stub that returns a fixed sequence of responses,” and adds that “Exact-match belongs to the deterministic layer.” Rows 1, 6 and 9 live there, and a rerun gives the same answer every time.
For a live model, ask for repeated runs, per-task pass rates and one pooled rate with its interval. Michigan requires an “Eval suite: at least eight tasks, each run at least three times, reported as pass rates,” which is 24 runs. Per-task intervals on three runs are too wide to compare, so the tolerance applies to the pooled rate.
How far apart may two honest numbers be?
Two honest pooled rates from 24 runs each can differ by about 25 points, so that is the tolerance I would state. The 95 percent margin for a difference of two proportions is 1.96 × √(2 × p × (1 − p) / n). At p = 0.75 and n = 24 that is 1.96 × √(2 × 0.1875 / 24) = 1.96 × 0.125 = 0.245, or 6 of 24 runs (illustrative).
In my simulation with independent runs, an honest student lands within 6 runs of the grader 94 to 99 percent of the time for true rates between 50 and 90 percent. When the grader reruns the same tasks, differences between tasks make two honest pooled rates land closer together than this, so the margin is generous on that count. What widens the gap is a rerun on a different setup: another machine, another build of the model, other sampling settings. Say in the handout that a miss triggers a second rerun and costs the points for that criterion only if both miss.
This tolerance catches a number that is far off: a reported 18 of 24 is accepted under 1 percent of the time when the agent’s true rate is 25 percent. It does not catch an invented number near the truth, since the same claim passes 58 percent of the time at a true rate of 50 percent. The submitted run records and the per-task pattern do more against that than any margin.
For the interval, use the eval sample size calculator, and the pass@k calculator shows what a per-attempt rate implies for passing every one of k runs. The neighbor post’s rubric says the rerun “matches”; that holds for a scripted client, and this tolerance is for live models.
Worked example: what does the judge-calibration check look like?
The check recomputes agreement from the student’s own label table, and it exists because the headline number misleads. Take row 7 with an illustrative submission: 30 outputs, of which the student labeled 21 pass and 9 fail. The LLM-as-a-judge agreed on 19 of the passes and 6 of the fails, passed 3 outputs the student had failed and failed 2 the student had passed.
Observed agreement is 25 of 30, or 83 percent. Agreement expected by chance is (21 × 22 + 9 × 8) / 900 = 0.59, so Cohen’s kappa is (0.83 − 0.59) / (1 − 0.59) = 0.59. Now take a judge that answers “pass” every time. It agrees on 21 of 30, which is 70 percent, and its chance agreement is also 0.70, so its kappa is 0.
Chapter 16 describes this trap: “if seventy percent of your outputs are good and the judge answers ‘good’ every single time, it scores seventy percent while carrying zero information.” The tool below opens on the example, with the student’s labels as rows.
With JavaScript on, the LLM Judge Agreement Calculator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the LLM Judge Agreement Calculator on its own page to share a result by link.
It reports kappa 0.59 with an approximate 95 percent interval of 0.26 to 0.92. Thirty labels meet the chapter’s “a few dozen” and still leave an interval that wide, so grade the analysis of the five disagreements and not the kappa. The companion post on whether an LLM judge is reliable covers the biases a student should test for.
What does a rubric for one of these assignments look like?
A rubric for these assignments names the evidence, names the check applied to each piece, and says what a failed check costs. The skeleton below is my own; fill the brackets and delete the lines an assignment does not use.
ASSIGNMENT: [name] Pairs with: [chapter, lab]
Individual or team: [ ] Staff time budgeted: [N] minutes per student
Where: [take-home | on paper in class | in the room on (date)]
Due: [date | round one (date), round two (date)]
Your input: [your own agent | the supplied agent, same for everyone |
pack number ___ | an anonymized classmate's pack]
Model access: [scripted client | a model you serve | course endpoint | none]
WHAT YOU SUBMIT
1. Runnable part: [code | tests | eval suite | clean and planted copies
with sealed tests | none]. It runs with [command] and needs no
account we do not provide.
2. Run record: [transcripts | traces | recordings | none], unedited,
[N] runs per task for anything that calls a live model.
3. Numbers: [pass rates per task, and a pooled rate with its interval |
label table and agreement statistics | cost per run | none], each
with the command that produced it.
4. Written part: [criterion | diagnosis | error analysis | decision
memo | none], at most [N] words.
5. AI-use record: the tools you used and what for. Required. No points.
HOW WE CHECK IT (delete what does not apply)
A. Rerun: we run part 1 against [hidden defects | hidden agent variants |
hidden tests | your own two copies | a classmate's pack | our copy].
Deterministic results must match exactly. A pooled rate from a live
model must land within [M] points of ours. If it does not, we rerun
once more; if it misses again, you lose the points for that
criterion only.
B. Read: we compare part 4 with [a key you have not seen | a second
reader's verdicts | the failures in our rerun].
C. Talk: [N] minutes, [in person | by video], between [dates]. We may
ask you to [change one thing | relabel K items | answer a what-if].
[Everyone | a random (fraction) of the class] is called.
CRITERIA Check Points
[criterion 1] [A|B|C] [ ]
[criterion 2] [A|B|C] [ ]
[criterion 3] [A|B|C] [ ]
LEVELS, for each criterion
Full: [the check passes: what we observe]
Partial: [what we observe, and the share of points]
None: [what we observe]
EARNS NO CREDIT
- A result about a live model taken from a single run.
- A failure of the infrastructure counted as a failure of the agent.
- (Only if check C is used.) A part of the submission you cannot
explain in check C or in the second conversation the course policy
allows.
AI USE
Allowed for: [ ]. Not allowed for: [ ].
Using a tool does not lower a grade. Part 5 records it; it is not graded.
Two lines are borrowed. Michigan’s rubric says that “Model disclosure and an evidence index are required but carry no points of their own: they are what makes the rest of the rubric checkable.” Keeping infrastructure failures out of the quality number is Chapter 16’s rule for any eval set.
What do the reported oral formats cost?
The four formats I can cite cost ten to twenty staff-minutes per student, and none reports outcomes at scale. Two take ten minutes per student one-on-one and one takes twenty. The group format used three staff for ten half-hour sessions to examine 58 students, which is 10 × 30 × 3 = 900 staff-minutes, about 15.5 per student (my arithmetic). Grouping shortens the calendar, not the staff time.
| Format | Who reports it | Time | The limit the source states or shows |
|---|---|---|---|
| One-on-one defense, comments stripped from the code beforehand | Engelhardt (2026), physics, Francis Marion University | At most ten minutes per student | One semester of an informal version in two courses “with enrollments of three and five”; the stripped-comment protocol and rubric were untested when he wrote |
| Individual, closed-book quiz after each homework | Stanford CS329Z, Fall 2026 syllabus | Ten minutes, twice a term | The page does not say how the quiz is staffed, and a syllabus reports no outcome |
| Group conversational exam with live coding | Barba and Stegner (2026), George Washington University | 58 students in groups of five or six, ten half-hour sessions over two days, “a team of three” | One examination, described by the two instructors who ran it |
| One question per student | A participant in the 2026 workshop | Twenty minutes per student | A participant’s account, unmeasured |
What to ask comes from the workshop report: “An oral assessment should not consist only of explaining a submitted artifact, because that can be rehearsed; it should require extending the code or making a change in real time, or at minimum asking what the student would do to accommodate a new requirement.” The ten-minute defense is from Engelhardt (2026).
The report also says where this stops, calling oral assessment “the strongest evidence of individual understanding” and “the least scalable thing anyone proposed.” A commenter who describes experience teaching and designing labs said it more bluntly in 2025: “Oral is such a time sink and needs so many TAs and lots of space if you want to run them in parallel” (Hacker News).
What will this cost to grade?
It costs hours once for every rerun, minutes per student for every read, and ten minutes or more per student for every live check, so the mix of actions sets the bill. The workshop report’s heading for this is “Scale is the binding constraint.”
Here is one illustrative term, with my estimates and the arithmetic shown. The class has ninety students, one instructor and two teaching assistants, and it uses rows 1, 5 and 7.
| Assignment | Once | Per student | Total for 90 |
|---|---|---|---|
| 1. Break your own loop (rerun) | 6 hours to write the broken copies and the script | 5 minutes each for a hand check of one submission in ten (9) | 6 h + 0.75 h = 6.75 h |
| 5. First wrong step (read) | 4 hours to record the pool and write the key | 6 minutes × 90 = 540 minutes | 4 h + 9 h = 13 h |
| 7. Calibrate a judge (rerun, read, talk) | 3 hours for the script, the pool of outputs and the ten keyed items | (5 read + 5 live + 2 changeover) minutes × 90 = 1,080 minutes | 3 h + 18 h = 21 h |
The three come to 40.75 staff-hours (6.75 + 13 + 21), about 13.6 hours each for three people across the term. Calling a random third of the class for the live part of row 7 makes it 3 h + (5 × 90 + 7 × 30) minutes = 3 h + 11 h = 14 h, and the total 33.75 hours.
At 30 students the same three cost 6.25 + 7 + 9 = 22.25 hours, and setup is more than half of it. At 300 they cost 8.5 + 34 + 63 = 105.5 hours, or about 82 with a random third on the live check (8.5 + 34 + 39.67).
A full ten-minute defense for everyone is (10 + 2) minutes per student: 6 hours at 30 students, 18 at 90 and 60 at 300, for a single assignment. Row 10, at 5 + 10 + 2 = 17 minutes per student, is 25.5 hours at 90. I would spend live minutes on two assignments a term.
Two costs sit outside the table. Rows 2 and 3 make the grader rerun every submission’s tasks three times on a live model, and row 2 once more per defective variant; that compute lands on the course’s machines or endpoint and grows with enrollment, so rerun a random sample when it does not fit. Student access is covered in the neighbor post’s section on whether students need a paid API; no row here requires a student to buy anything.
What should the syllabus say about AI tools?
It should allow the tools, say that the grade comes from what the staff can rerun, compare or ask about, and name where understanding is checked. That is my position. Two Fall 2026 syllabi share its first two parts: tools allowed, and you must be able to explain what you submit.
Stanford’s CS329Z: “This is a course about building with AI, so we expect you to use it,” with AI-generated code permitted “provided that you can validate and explain all code. This will be verified through oral examinations.” Michigan’s EECS 498: “AI use is required in this course; that is the point,” followed by “You must understand every line you submit and be able to explain it.”
They differ from my paragraph on one point. Stanford says that “using AI to substantially complete an assignment without actively understanding the work is an Honor Code violation,” and Michigan’s page says violations of its rules “go through the College of Engineering Honor Code process.” My paragraph treats an unexplained part as a grading matter, which your institution may not allow; the bracket in it tells you to check.
AI TOOLS IN [COURSE CODE]
This policy supplements [the institution's AI and academic integrity
policies]; where they differ, those govern.
You may use AI tools, including coding agents, for [all graded work |
all graded work except: ___]. This course is about building with models,
and using them well is part of what it teaches.
Your grade does not rest on the files alone, because a tool can produce
the files. It rests on what the staff can check: we rerun what you
submit, we compare your analysis with a key, and we may ask you about
your work for [N] minutes; [everyone | a random part of the class] is
called each time. Each handout says which of these applies.
You are responsible for everything you submit. A part of your submission
that you cannot explain when asked earns no credit, whoever or whatever
wrote it. If you cannot explain a part on the day, you may ask once for
a second conversation with [a different staff member] within [N] days.
[Instructor: before using this paragraph, check with your integrity
office that a zero on this basis may be given as a grading decision.]
Reporting a result, a run record or a label that was not produced as
described is fabrication and goes to [the institution's process].
With each submission, list the tools you used and what you used them
for. The list is required, carries no points and cannot lower your grade.
Not allowed: [copying another student's code, prompts or reports];
[using a tool to write ___]; [any tool during a live check, unless the
handout says otherwise].
Access: every assignment can be completed with [a scripted client |
a model you serve on your own or a lab machine | the course endpoint].
No assignment requires a paid account. If access is a problem, tell the
staff by week [N].
If a disability, a language need or a schedule conflict affects a live
check, contact [the accommodations office or name] by [date] and we
will arrange [another format or time].
The no-credit rule is mine, and it goes further than the workshop report does. For cases where misconduct cannot be established, the report’s advice is that “the workable response is structural rather than punitive: downweight the artifact, increase the weight on the oral assessment.” That reweights course components; my rule zeroes a part, which is why it carries a second attempt.
Does an assignment you already have still measure anything?
It does only if the answers to questions 1 and 2 below are yes; the other six tell you what the check will cost and where it is weak. The list is mine. Run it once for each assignment you plan to keep.
- Regrade. Could I regrade this next month from what I kept: the submitted files, by rerunning them or by reading them against a key, or a marking sheet from the live check?
- Unseen. Is at least one thing I grade against hidden from the student until grading: a key, a defect, a task or a question?
- Runs. Does every number about a live model come from at least three runs per task, with the runs submitted? (Tick it if no live model is involved.)
- Tolerance. Does the handout say how close my rerun must land to the student’s number? (If the assignment reports numbers and nothing is rerun or checked against a key, the answer is no.)
- Own input. Does each student work on something nobody else has: their own agent’s runs, or a draw from a pool?
- Watched. Does a grader or a proctored room see the student redo, change or extend part of the work, in this assignment or in the one it feeds?
- Minutes. Have I multiplied the staff minutes per student by the enrollment, and does the product fit the staff I have?
- Policy and access. Does the handout say what AI use is allowed, how it is recorded, and that the work can be done with access that costs the student nothing?
A build-and-demo assignment with a README fails on questions 1 and 2: nothing can be rerun against anything hidden, and no key exists for a demo. It can still answer yes to 5, 7 and 8, which is why a count of ticks is the wrong test. Row 8 passes question 1 only if the examiner fills in a marking sheet during the conversation; without one, the grade is the examiner’s memory.
Row 1 of the bank passes 1 and 2 and answers no to 5 and 6, which is the profile of a practice assignment. Row 9 passes 1, 2, 5 and 6, and its weak point is outside the list: the tool is in the room.
What does none of this fix?
No design here stops a determined student from having a model do the work and then studying the result until they can defend it. In the workshop report, one room “could not identify a reliable way to measure the capacity to do something without AI” short of invigilating a two-hour session. A course that needs that guarantee needs a supervised exam, and the teaching kit’s assessment split keeps one at 20 percent.
Nothing here is measured. I know of no outcome data showing that AI agents assignments for students resist model completion better when graded by rerun, and the bank’s proposals have not been run in a course I can cite. The course pages were read on 2026-10-06 and may change during the term.
The live checks do not scale past a certain enrollment, and the report raises its example of critique at “400-plus students and a handful of TAs” without solving it. They carry a fairness cost too: the report notes “the complications oral assessment introduces for international and disabled students,” and that an inability to explain a system “may reflect weak understanding, weak communication skills, or both.” Under sampling, only the students called can lose credit in the conversation.
Reruns grade the artifact, so a suite written by a model and never read can earn full marks. Chapter 16’s warning applies to a gradebook as much as to a deployment: “a green suite is evidence, never proof.” Hidden material leaks once the first cohort has seen it, so plan to refresh defects, variants and trace pools each term.
Where to go from here
The one idea to keep is that a grade is a claim, and it deserves the treatment the book asks for an agent’s claim that it has finished: find the signal you can check without trusting the claimant. For each assignment, write into the handout whether you will rerun it, read it against a key or talk about it, and weight it by what that check can tell you.
Students will ask why the course grades this way. My answer is that the best way to learn AI agents is to build one you can verify, and the grade follows the same rule.
The week-by-week plan, the published labs and the browser exercises are in the free teaching kit. Chapter 3, “The Agent Loop,” Chapter 15, “Observability and Debugging,” and Chapter 16, “Evaluating Agents,” are in the full book, and you can see the formats.
Questions readers ask
- If answers are no longer proof of work, what should students submit for an AI agents assignment?
- Ask for things the grader can verify without taking anyone's word: tests or an evaluation suite the grader reruns against defects the student has not seen, a diagnosis the grader reads against a key, and recorded runs with pass rates over at least three runs per task. Those checks grade the artifact, whoever wrote it. To learn something about the student, add a short live check or do part of the work in the room.
- How do we know a process log or a reflection reflects the real process and is not AI-generated?
- You cannot know from the text. In the report of a 2026 working-group workshop of computing educators, the participants in one room who collect process evidence said they do not grade it, and the report concludes that nobody in the workshop identified process evidence that resists being generated. Treat a log as a learning aid, and put the grade on things you can rerun, compare with a key, or ask the student about.
- How do oral checks scale to classes of 30, 100, or more?
- They scale by sampling, and only so far. Grouping shortens the calendar, not the staff time: one published group format used three staff for ten half-hour sessions to examine 58 students, about 15 staff-minutes per student. For ninety students a ten-minute check with two minutes of changeover is 18 staff-hours per assignment (an illustrative calculation), so a course can check a random third each time or keep live checks for two assignments a term.
- How do we grade fairly when not every student has paid AI access?
- Design every assignment so it can be completed with access that costs the student nothing: a scripted model client for harness work, a model the student serves on a laptop or a lab machine, or a course-provided endpoint for sessions in the room. Then grade things that depend less on the tool: whether numbers reproduce, and what the student can do in the room. A stronger tool still writes a stronger test file, so keep the weight of take-home work low.
- Should students be allowed to use AI tools in an AI agents course?
- The two build-centered syllabi quoted in this post, Stanford's CS329Z and Michigan's EECS 498 (both Fall 2026), expect or require it and verify understanding in person. This post takes the same position on those two points: allow the tools, grade what the staff can verify, and say so in the syllabus.
Sources
- Akbar, Challen, Fund, Hopkins, Karnalim, Lin, McGuffee, Taneja and Ware (SIGCSE Virtual working group) (2026). AI Can Do Your Homework. Now What? Report from an Online Workshop on Computing Assessment in the Age of Generative AI
- Stanford University (2026). CS329Z Engineering AI Agents: Logistics (Fall 2026)
- University of Michigan EECS (2026). EECS 498-016 Applied Agentic Software Engineering: Syllabus (Fall 2026)
- Stanford University (2025). CS329T Trustworthy Machine Learning (Fall 2025), Homework 1: Measuring and Improving Agent GPA
- Larry Engelhardt (2026). A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI
- Lorena A. Barba and Laura Stegner (2026). The Conversational Exam: A Scalable Assessment Design for the AI Era
- Yusuf Pisan (2026). Teaching Intro AI When the Tools Can Do the Homework: A Course Redesign and a Student Bill of Rights
- kingstnap (2025). Hacker News comment on the cost of oral assessment