Repository navigation
Programming Question Snapshotting / Audit Trail - #8611
Conversation
…ing services and job - resolve rubocop todos in the grading services Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Import and grading concurrency windows can still mismatch packages, remote problems, snapshots, and pinned test cases.
Review effort: Balanced
Findings: 3
Open (3)
What changed in this PR
Adds immutable programming-question snapshots and preserves grading history across edits and regrades.
Changes:
- Snapshots replaced question versions and retain test cases, templates, and package references.
- Stores multiple programming grading runs and narrows regrading to visible/current answers.
- Displays historical results against their original tests with a version-change notice.
| File | Description |
|---|---|
spec/services/course/assessment/question/programming_import_service_spec.rb |
Tests snapshot imports and rollback behavior. |
spec/services/course/assessment/question/answers_evaluation_service_spec.rb |
Tests narrowed regrading scope. |
spec/services/course/assessment/answer/text_response_comprehension_auto_grading_service_spec.rb |
Updates grading API tests. |
spec/services/course/assessment/answer/text_response_auto_grading_service_spec.rb |
Updates grading API tests. |
spec/services/course/assessment/answer/rubric_auto_grading_service_spec.rb |
Updates grading API tests. |
spec/services/course/assessment/answer/programming_codaveri_auto_grading_service_spec.rb |
Updates Codaveri grading tests. |
spec/services/course/assessment/answer/programming_auto_grading_service_spec.rb |
Tests version pinning and retained runs. |
spec/services/course/assessment/answer/multiple_response_auto_grading_service_spec.rb |
Updates grading API tests. |
spec/services/course/assessment/answer/auto_grading_service_spec.rb |
Tests explicit grading runs. |
spec/models/course/assessment/question/programming_spec.rb |
Tests snapshots, immutability, and package removal. |
spec/models/course/assessment/answer/programming_spec.rb |
Tests completed-run selection. |
spec/models/course/assessment/answer_spec.rb |
Tests multi-run associations and creation. |
spec/jobs/course/assessment/answer/reduce_priority_auto_grading_job_spec.rb |
Updates job argument tests. |
spec/jobs/course/assessment/answer/auto_grading_job_spec.rb |
Tests explicit run persistence. |
spec/helpers/course/assessment/answer/programming_test_case_helper_spec.rb |
Tests grading-version resolution. |
spec/controllers/course/assessment/submission/answer/answers_controller_spec.rb |
Tests historical-result rendering. |
spec/controllers/course/assessment/question/programming_controller_spec.rb |
Tests package removal and snapshot isolation. |
lib/tasks/db/migrate_programming_question_languages.rake |
Excludes snapshots from language migration. |
db/schema.rb |
Records snapshot columns and non-unique run index. |
db/migrate/20261006000001_drop_unique_answer_id_index_from_auto_gradings.rb |
Concurrently removes the unique run index. |
db/migrate/20261006000000_add_non_unique_answer_id_index_to_auto_gradings.rb |
Concurrently creates its replacement index. |
db/migrate/20261002000000_add_snapshot_columns_to_course_assessment_question_programming.rb |
Adds snapshot metadata and foreign keys. |
client/locales/zh.json |
Adds the historical-version notice. |
client/locales/ko.json |
Adds the historical-version notice. |
client/locales/en.json |
Adds the historical-version notice. |
client/app/types/course/statistics/answer.ts |
Types the historical-version flag. |
client/app/types/course/assessment/submission/answer/programming.ts |
Extends programming-answer state. |
client/app/bundles/course/assessment/submission/reducers/gradingResults.ts |
Carries the version flag into state. |
client/app/bundles/course/assessment/submission/reducers/__test__/gradingResults.test.ts |
Tests version-flag state updates. |
client/app/bundles/course/assessment/submission/containers/TestCaseView/index.tsx |
Passes the flag to test-case rendering. |
client/app/bundles/course/assessment/submission/components/AnswerDetails/ProgrammingComponent/translations.ts |
Defines the notice message. |
client/app/bundles/course/assessment/submission/components/AnswerDetails/ProgrammingComponent/TestCases.tsx |
Renders the historical-version alert. |
client/app/bundles/course/assessment/submission/components/AnswerDetails/ProgrammingComponent/__test__/TestCases.test.tsx |
Tests alert visibility. |
client/app/bundles/course/assessment/submission/components/AnswerDetails/ProgrammingAnswerDetails.tsx |
Propagates the version flag. |
app/views/course/assessment/answer/programming/_programming.json.jbuilder |
Serializes historical tests and results. |
app/services/course/assessment/question/programming/programming_package_service.rb |
Routes package removal through snapshotting. |
app/services/course/assessment/question/programming_import_service.rb |
Creates snapshots during imports. |
app/services/course/assessment/question/answers_evaluation_service.rb |
Regrades only visible/current answers. |
app/services/course/assessment/answer/text_response_comprehension_auto_grading_service.rb |
Writes into an explicit grading run. |
app/services/course/assessment/answer/text_response_auto_grading_service.rb |
Writes into an explicit grading run. |
app/services/course/assessment/answer/rubric_auto_grading_service.rb |
Writes into an explicit grading run. |
app/services/course/assessment/answer/programming_codaveri_auto_grading_service.rb |
Pins tests and preserves Codaveri runs. |
app/services/course/assessment/answer/programming_auto_grading_service.rb |
Pins tests and preserves programming runs. |
app/services/course/assessment/answer/multiple_response_auto_grading_service.rb |
Writes into an explicit grading run. |
app/services/course/assessment/answer/auto_grading_service.rb |
Passes and saves explicit runs. |
app/models/course/assessment/question/programming.rb |
Integrates snapshots and package removal. |
app/models/course/assessment/question/programming_test_case.rb |
Protects snapshot test cases. |
app/models/course/assessment/question/programming_template_file.rb |
Protects snapshot templates. |
app/models/course/assessment/answer/programming.rb |
Selects the latest completed run. |
app/models/course/assessment/answer/auto_grading.rb |
Allows multiple runs per answer. |
app/models/course/assessment/answer.rb |
Adds run history and run creation logic. |
app/models/concerns/course/assessment/question/programming_snapshots_concern.rb |
Implements snapshot creation and associations. |
app/models/concerns/course/assessment/question/programming_snapshot_read_only_concern.rb |
Enforces snapshot immutability. |
app/jobs/course/assessment/question/programming_import_job.rb |
Carries pre-edit values into imports. |
app/jobs/course/assessment/answer/base_auto_grading_job.rb |
Carries the selected grading run. |
app/helpers/course/assessment/answer/programming_test_case_helper.rb |
Resolves the graded question version. |
.rubocop_todo.yml |
Removes resolved exclusions. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
- track snapshot creator (superseder) and creation time (superseded_at) - fix validate_language_enabled to be within the programming question model
- added alert for past answer grading results done before edit
- convert ProgrammingAutoGrading has_one association into has_many
41c27a0 to
de9eed8
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Import and grading races can still corrupt or misattribute preserved grading history.
Review effort: Balanced
Findings: 5
Open (6)
Skip superseded jobs before applying imports · New Scope Codaveri bulk updates to live records · New Atomically record jobs and handle enqueue failures · New Read package data and tests from one consistent version · New Pin evaluation results to the Codaveri-evaluated version · New Sync the current question version before Codaveri evaluation · New
…ing logic - make bulk codaveri-related updates only affect live questions
de9eed8 to
39f891c
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved issues affect version consistency, Codaveri synchronization, and visible grading results.
Review effort: Lite
Findings: 4
Open (4)
Resolved since last review (6)
Pin evaluation results to the Codaveri-evaluated version Read package data and tests from one consistent version Atomically record jobs and handle enqueue failures Scope Codaveri bulk updates to live records Skip superseded jobs before applying imports Sync the current question version before Codaveri evaluation


Summary
Editing a programming question today destroys the record of how its existing answers were graded, then
regrades every answer ever submitted against the new version. This PR keeps that record:
snapshot instead of being deleted, so every past test result still points at the test case that produced
it.
were given.
previous one.
has changed since.
same question can no longer interleave, and an import that finishes after a later edit has committed is
dropped rather than overwriting it.
This replaces the draft in #8022 (and its predecessor #7595), rebuilt on two changes that have since merged:
the test case payload split and TypeScript test case views (
f93da269e), and instance admin assessmentpermissions (
9d461ccd7). Nothing below assumes familiarity with the earlier drafts.It is also a prerequisite for the programming language upgrade feature, whose course-wide re-import was
judged too risky to run while every re-import destroyed grading history.
The problem
When a programming question's package is re-imported (a new package, or a change to its language, time
limit or memory limit),
ProgrammingImportService#save!assignsquestion.test_cases = <new records>.There is no diffing: every test case row is destroyed and recreated, even when nothing in it changed. Test
results are
dependent: :destroyfrom their test case, so every per-test-case result of every answer tothat question goes with them.
AnswersEvaluationServicethen regrades every non-attempting answer eversubmitted.
So after any edit:
What changes, at a glance
test_case_idstill resolvesCommits
Each commit can be reviewed on its own.
refactor(auto-grading): pass the auto grading record through the grading services and jobanswer.auto_grade!obtains theAutoGradingrecord (existing or new) and passes it explicitly through thejob and every grading service, instead of each layer rediscovering it with
answer.auto_grading. This is ano-op on its own, and is what makes "one run per grading" possible later. It changes the grading jobs'
arguments (see Deployment).
feat(programming): create snapshots of programming questions on editSnapshot creation and the snapshot columns migration. Also how imports and grading runs coexist:
feat(programming): only regrade current/latest answers on question editThe narrowed regrade, and rendering old results against the test cases they were graded with.
feat(programming): make question snapshot rows read-onlyThe immutability guard, the snapshot on package removal (which also fixes the 500 above), and specs
showing snapshots are unreachable by URL.
feat(programming): preserve previous auto grading runshas_many :auto_gradings, and the two migrations that swap the uniqueanswer_idindex for a non-uniqueone without locking.
feat(programming): handle concurrent imports / improve codaveri handling logicimport" well defined.
Every commit also clears the
.rubocop_todo.ymlentries for the files it touches, where that was possiblewithout unrelated refactoring.
Design
What a snapshot is
A snapshot is a row in the programming question table with
current_idpointing at the live question. Ithas no parent
course_assessment_questionsrow, so it is never part of an assessment, never listed,and never reachable through any question query. It holds:
one reference row and no storage;
superseded_at(when the import replaced it) andsuperseder_id(whose edit did). The original story istitled "Programming Question Audit Trail", and these are the minimum an audit trail needs.
Things a reviewer may wonder about:
dup? Underacts_as,dupduplicates the parent question too. Every edit would leave anorphaned parent row carrying the new title and maximum grade, visible to every global question query and
not deleted with the snapshot. The snapshot row is inserted directly (
insert!), which also skips thecallbacks and validations meant for an editable question.
values in the request; the import that creates the snapshot runs later, in a job. By then the row already
holds the new values, so the snapshot would pair the new language with the old test cases, a combination
that never existed.
Programming#process_package(before_save) captures the values still in the databaseand passes them to
ProgrammingImportJobas a new, optional argument. They include the editor, taken fromUser.stamper, because the job runs without a user.column_names, notattribute_names. Underacts_asthe latter includes the parent question'scolumns.
Where the snapshot is taken, and why it matters
The snapshot is taken and the old test cases are moved inside
ProgrammingImportService#save!, in thesame transaction that inserts the new test cases. Not in the controller, and not as a separate step.
The reason is a full-marks window. A programming question with no test cases is not auto-gradable, and
grading such a question awards full marks without evaluating anything. That check runs against the live
question at the moment of grading. If the controller moved the old test cases away when the edit was saved,
the live question would have none until the import job finished evaluating the new package, which takes
seconds to minutes and never finishes if the import fails. Every answer graded in that window would get full
marks, published under the System grader in autograded assessments. The narrowed regrade would not correct
earlier attempts graded in that window.
Today's code has no such window: replacing a saved association runs its deletes and inserts in one
transaction, so other connections see the old test cases or the new ones, never none. Doing the move in the
same transaction keeps that property. It also means a failed import leaves no snapshot and changes nothing,
exactly as today.
The move is
question.test_cases.update_all(question_id: snapshot.id), which is O(test cases) and copies noresults. It must go through the association:
update_allon an association resets it, whereas throughthe model class a preloaded association still holds the moved rows, and the next assignment deletes them. A
spec with preloaded associations covers this.
A grading run in flight during an edit
Both programming graders spend seconds to minutes evaluating before they match results to test cases, so an
import can commit in the middle of a run.
Local evaluator. A run reads the question's test cases and its package in one
REPEATABLE READtransaction (
pin_version): a single database snapshot. Every change to a question's package changes thepackage and test cases together, in one transaction. So the run gets either the old version or the new one,
never one version's package with the other's test cases.
answer shows the "changes have been made" alert until the edit's regrade adds a run against the new
version.
nil.A spec commits an import from a second database connection between the two reads; without the shared
snapshot, the run evaluates the new package against the old test cases. The other two windows are covered
too:
Codaveri. Nothing local can be pinned, because the problem being evaluated lives on Codaveri and an import
replaces it. A pin wouldn't have survived anyway: the pre-evaluation sync's
with_lockreloads the question,which unloads its test cases. Instead, the run is attributed to the version Codaveri actually evaluated. Each
result names the test case it is for (its
indexis the local test case id pushed to Codaveri), and snapshotskeep every version's test cases. The run's counts and failure records come from that version. Results that
don't all belong to one version of this question raise a
CodaveriErrorsaying so; today the same situationcrashes with
NoMethodErroronnil.Which version a run was graded against: derived, not stored
Earlier drafts stamped a
question_snapshot_idon each grading run. That column would sit on a 72 GBtable, and it goes stale: a run stamped with the live question becomes wrong the moment an edit moves its
test cases to a snapshot. Keeping it correct meant a mass
UPDATEon every edit.Instead, the version is read from the run's own results:
test_result → test_case → question_idis the livequestion or one of its snapshots. A test case is never deleted on edit, and a snapshot never changes, so this
cannot go stale. It covers every grading path that writes
test_case_id, including Koditsu's bulk insert,and both columns involved are already indexed.
One legacy case: a run whose results were deleted by an edit made before this PR has a grading record but no
results, so it says nothing about its version. It is treated as "graded on an earlier version" and rendered
against the current test cases with no marks, as today, plus the alert. A completed run is never otherwise
empty, because the grader records a failure for every test case the report misses.
Which answers are regraded after an edit
Only answers whose grading is shown outside a submission's Past Answers. Per submission, these are:
attempting submission (
SubmissionsHelper#last_attempt), and the "last attempt" in assessment statistics;shows once the submission is no longer being attempted.
Normally these are the same answer. They differ in submissions from before
d4294b5c7(2021-11-25). Beforethat commit, finalising marked the current answer submitted in place, so it could be older than attempts
submitted after it. Since then, finalising promotes the last submitted attempt to current, or finalises a
fresh copy. In a production-scale copy, about 3.0M submitted current answers (one per question per
submission) created before that change are older than a submitted attempt at the same question. So are 10,732
created after it, whose cause was not pinned down. By submission state, those are 10,537 published, 99
submitted, 92 attempting and 4 graded. For these, the edit page and the statistics show different answers, so
both are regraded.
Answers still being edited are never regraded. Autosave only saves the answer's files; it neither grades nor
changes
created_at.What the UI shows
Every view continues to render the current question, including title, description and highlighting
language. The one exception is the test case panel, which renders the test cases of the version the run
was graded against, so that its pass/fail marks, hints and evaluation check make sense. When that version is
not the live one, the panel shows an alert: "Changes have been made to the question after this answer was
graded." This applies on the submission edit page, in Past Answers and in assessment statistics.
Grading runs per answer
Answernowhas_many :auto_gradings(oldest first). The existinghas_one :auto_gradingis kept as thelatest run, ordered newest first, so every existing reader keeps working unchanged. Runs are only
created through
auto_gradings.in place, as today: they have nothing worth comparing, and this keeps their regrades from growing a 35M-row
table.
results, with the "being evaluated" or error status from the latest run. This matches what a re-evaluate
shows today, where the old results stay attached until replaced. After an edit, it is an improvement: today
the edit has already destroyed the old results. Without it, the panel would empty for the whole
low-priority regrade after an edit.
resubmit_programming, when anassessment's graded test case types change) finalises and regrades immediately, so it now just adds a run.
ensure_auto_grading!relied on the unique index and is gone with it. The worstcase is that two concurrent first gradings of a non-programming answer each create a run, and the later one
is used.
No comparison UI is built here. The data model now supports one.
Overlapping edits and imports
Package evaluation takes seconds to minutes, so imports of one question can overlap:
Two problems follow.
snapshot. The second import's move waits on the first's row locks. Once the first commits, Postgres
re-checks the second's condition, finds the rows already moved, and moves nothing. So the second
snapshot is empty, and the second import's assignment of new test cases then destroys the first
import's, with any results graded against them. A spec running two imports on separate connections
reproduces exactly this when the lock below is removed.
the order of the edits. An older edit's import finishing late overwrites the newer edit's package, while
the question's settings stay the newer edit's.
The fix.
Codaveri push all take the same
SELECT … FOR UPDATElock on the question's row first (lock_package!,or
with_lockfor the push). The import then re-reads everything under it: the existence check, theprevious package and the test cases. The lock is held only for that short write, never during evaluation.
The local grader never takes it; Codaveri grading takes it only for the sync it runs before evaluating.
id to
import_job_idin the same transaction as its other changes, and enqueues the job only after commit.Before, the id was written after commit and after enqueueing. A job could then start before its own id was
recorded, and two concurrent edits could record their ids in the opposite order to their commits. Edits
that schedule an import all update the question's row, so Postgres orders them, and the recorded job is
always the latest committed edit's.
skip a package that won't be applied, and again under the lock, which is the check that decides. A
superseded import changes nothing, queues no regrade, pushes nothing, and its job completes normally.
Removing the package clears the recorded job, so it supersedes any import still pending.
import_job_id, for the edit page to follow. They now do so only when no import is pending; otherwisethey would make that import look superseded.
question waiting forever on an import that was never queued.
What is guaranteed. The package, test cases and template files always equal exactly one applied change:
the original state, one import's output, or a package removal, never a mix. History has no gaps: each import
snapshots the version it actually replaced. Once a later edit has committed, no earlier import can apply. An
earlier import that finishes before the later edit commits does apply, and is kept as a snapshot when the
later one applies.
Whether the final state matches the latest edit depends on that edit's import:
submitted)A snapshot pairs the package that was live with the settings that were live just before it was replaced.
So if a time-limit edit's import was superseded, the snapshot shows the new limit with the old package.
That's accurate: it is what grading used in that window.
Keeping Codaveri in step
Codaveri holds its own copy of each Codaveri question, pushed from the local one.
safe_create_or_update_codaveri_question. That takes the row lock and sends the question's currentpackage and test cases. Before, the import job and
CodaveriImportJobpushed the package they were queuedwith, so a late push could leave Codaveri on an older version than the local question.
Codaveri lags behind; gradings in that window are attributed to whichever version Codaveri evaluated
(above).
failure status written to the question was rolled back with it. It now re-raises after the transaction
commits, so the status is kept, as the import job's direct push always kept it.
Snapshots are read-only and unreachable
ProgrammingSnapshotReadOnlyConcernraisesActiveRecord::ReadOnlyRecordon saving a snapshot, or a testcase or template file that belongs to one, including moving one onto or off a snapshot. It runs before
validation, so the refusal is explicit rather than an unrelated validation error. Destroying is allowed, so
that deleting a question deletes its snapshots. Callbackless writes (
insert!,update_all) bypass itdeliberately: they are how snapshots are created.
ProgrammingControlleraction loads its question through@assessment.programming_questions, whichrequires the parent question row a snapshot lacks. So the edit URL, and every other action, return 404
for a snapshot id. If that loading ever changed, the ability check would still refuse (
AccessDenied). Aread-only snapshot view is left for the comparison UI.
Lifecycle
reachable from an assessment, so they are not duplicated.
has_many :snapshots, dependent: :destroy), with their testcases and package references.
new
Programming#remove_package. This also fixes a bug:test_cases.clearissues a rawDELETEthatskips callbacks, which the results' foreign key rejects, so the switch failed with a 500 once any answer had
been graded.
db:migrate_programming_question_languages) now changes live questionsonly, so snapshots keep the language they were graded under.
Smaller fixes along the way
validate_language_enabledwas defined afterProgramming's closingend, which made it a private methodon
Object. It is now a private method of the class, with model specs.acts_asgem overridesexists?with an inner join to the parent question row, soProgramming.exists?(snapshot_id)is always false, andsnapshots.empty?on an unloaded association isalways true. Specs that relied on these were passing vacuously; they now use
ids/pick, and the guardavoids
exists?.Deployment
Migrations
20261002000000_add_snapshot_columns_…current_id,superseded_atandsuperseder_idto the 235k-row programming question table (16 MB): nullable, no default, no backfill, partial indexesWHERE … IS NOT NULL, foreign keys20261006000000_add_non_unique_answer_id_index_to_auto_gradingsidx_course_assessment_answer_auto_gradings_on_answer_idconcurrently on the 35M-row auto gradings table20261006000001_drop_unique_answer_id_index_from_auto_gradingsanswer_idindex concurrently, after checking that M2's index is validMeasured locally on a production-scale copy (35.1M auto grading rows, unique index 752 MB):
During M2 and M3, a probe wrote to an auto grading row twice a second, each in a rolled-back transaction.
Write latency stayed at baseline (mean 40 ms during M2 against 55 ms before, max 75 ms, mostly
psqlstart-up), with no errors. After M3, the latest-run lookup (
WHERE answer_id = ? ORDER BY created_at DESC, id DESC LIMIT 1) uses the new index (0.2 ms). These are laptop numbers against an idle database; stagingwill be the realistic check.
M2 and M3 are separate migrations so that Rails records each one on its own. These are the repo's first
disable_ddl_transaction!migrations, so worth confirming that the deploy tooling tolerates them.How they fail, and how to recover:
CREATE INDEX CONCURRENTLY(timeout, cancelled deploy, lost connection, out of disk)leaves an INVALID index behind.
IF NOT EXISTSwould accept it as done, so M2 does not use it: it dropsan invalid leftover and builds again. Re-running M2 is always the recovery.
35M rows.
second run per answer would fail. Re-running M3 finishes the drop.
has a second run. Its
downraisesIrreversibleMigration. A code rollback does not need a schemarollback. The old
has_onewould read an arbitrary one of an answer's runs, which is degraded but notbroken.
All of these paths were exercised on the test database.
Order and preconditions
Deploy in a window with no autograding jobs queued or running. Commit 1 changes the grading jobs'
arguments, and jobs serialised by the old code would not deserialise correctly. There is deliberately no
compatibility shim.
Before M2:
pg_stat_activityfor long-running transactions, read-only ones included (reports, backups,stuck jobs), because a concurrent build waits for them;
Long-running transactions. The build waits for every transaction that is open when it starts, so anything
old here will hold it up:
Timeouts, as seen by the connection the migration actually uses. This includes anything set in
database.ymlor at the role or database level:Each should be
0, meaning disabled, or comfortably longer than the build. Role- and database-leveloverrides, if you need to find where a value comes from:
Disk. The new index is about the size of the one it replaces:
Compare that with free space on the database volume. On a self-hosted server, run
df -h "$(psql -Atc 'SHOW data_directory')"on the host; that needs superuser. On a managed database, usethe provider's storage metrics.
Run M1, M2 and M3, in order, before the new code serves traffic. M3 must complete first:
Running them one at a time shows each one's duration; a plain
bin/rails db:migrateruns all three in thesame order. While M2 builds, its progress is visible from another session.
current_locker_pidis thetransaction it is waiting for, if any:
Afterwards, check that the new index is valid and the unique one is gone. The query is in the staging
section below.
Deploy the code.
Old code that keeps serving between M3 and the code deploy has lost its database-level guard against two
concurrent first gradings of one answer each creating a run. Within the no-jobs window this is unlikely, and at
worst leaves a spare run.
Import jobs queued by the old code still run:
ProgrammingImportJobgains an argument, but it is optional and appended.recorded one and applies.
CodaveriImportJobkeeps its arguments; it now ignores the package it was queued with.Known gaps and follow-ups
language and limits in the request, while the package and test cases only change if the import succeeds. A
failed import therefore leaves the new settings paired with the last applied package and tests (see the
table under "Overlapping edits and imports"). Snapshotting does not make this worse, since a failed import
creates no snapshot. Making version changes all-or-nothing would change what staff see when an import
fails, and touches the flow the language upgrade feature relies on.
submitted) isn't retried automatically, andCodaveri-only edits won't record their sync jobs while it's pending. Only a failed import (
errored)triggers a retry on the next save.
its job, and only
Programming#schedule_importrecords it. AProgrammingImportJobenqueued directly isalways superseded. Nothing in this repository does that, but the language upgrade branch needs checking
before it merges.
snapshot and a regrade. The package's SHA-256 makes a cheap comparison possible.
needs care with in-flight runs, and isn't needed for correctness.
check_zombie_jobsregrades into a new run, so the dead run stays inthe history with its stuck job. A comparison UI should expect unfinished runs.
Programmingis within about 6 lines of Rubocop's 200-line class limit, even with the scheduling codeextracted into
ProgrammingImportsConcern. The next addition will need something else extracted first.create(:submission, :submitted, auto_grade: false)raisesNameErroron anundefined
answer.Testing
Automated
Backend: assessment models, services, jobs, controllers and helpers, plus duplication services: 1679
examples, 0 failures. New specs cover:
and superseder;
the run's two reads;
blocked. Without the lock, this reproduces the empty snapshot;
removal, and a superseded job queueing no regrade;
rejected push keeping its failure status,
CodaveriImportJobpushing the current package, and aCodaveri-only edit not displacing a pending import;
answers#showfor an answer graded before an edit, and while a later run is still grading;PATCH #update;Each was checked to fail with the behaviour it covers removed.
Frontend: full Jest suite passes; new tests for the alert and the
gradingResultsreducer carrying theflag.
Rubocop clean repo-wide.
One earlier full run had 2 failures in specs this change doesn't touch: reference-timeline duplication,
and the assessment controller's tabbed-view update. Both pass in isolation and in two later full runs with
different random orders. The cause wasn't found, and the new specs that start threads can't be ruled out;
worth watching in CI.
Not run: feature (browser) specs.
spec/features/course/assessment/submission/autograded_spec.rbis theone that touches grading runs.
Suggested test paths on staging
Run the migrations first, recording each one's time and checking the index state afterwards:
Use an assessment with a zip-upload programming question and a few students. One submission should be
attempting with several submitted attempts, and another should be submitted.
These console helpers make the checks quicker:
a.auto_gradings)resubmit_programmingis_synced_with_codaveriis true and the grading results match the new test casesimport_job_idstays the import's until it finishes. The live-feedback sync still runs and pushes the current version🤖 Generated with Claude Code