/hard-cheese
When to invoke: Metacognitive vibecheck gate before code is shared for review — make the author explain the diff’s causal logic, graded by a fresh-context judge against the SOLO Taxonomy. Use when the user wants this gate — phrases like “/hard-cheese”, “/cheese –hard”, “gate this before I push”, “vibecheck me”, “make sure I understand this diff”, “epistemic-debt check”. Use standalone before opening a PR, or as the --hard flag propagated through the pipeline. Do NOT use for code review (/age), test hardening (/press), or fix application (/cure).
The gate mitigates epistemic debt — the failure mode where AI-scaffolded code passes review, type-checks, and tests green while the author cannot explain it to a reviewer.
Inputs
Section titled “Inputs”/hard-cheese [<slug>] [--socratic-cap N=3] [--passing-score N=3] [--no-judge]Arguments:
<slug>— optional. Identifies the artifact at.cheese/hard-cheese/<slug>.md. When omitted, fall back to the git short SHA ofHEAD. An explicit slug always wins.--socratic-cap N— max retry attempts before the gate marks the artifactFAILEDand exits non-zero. Default3. Vibecheck does not cap; easy-cheese does to avoid infinite loops.--passing-score N— minimum SOLO score that counts as PASS. Valid range1..5; default3(Multistructural-or-higher). A previous PASS below the requested threshold is treated as stale and must be re-judged.--no-judge— log-only mode. Capture the user’s explanation, write the artifact withstatus: LOGGED, skip the judge sub-agent spawn. Mirrors vibecheck’s optional JSONL telemetry mode.
Invocation modes
Section titled “Invocation modes”| Mode | How it fires | Where the gate sits |
|---|---|---|
| standalone | User runs /hard-cheese <slug> directly before opening a pull request. |
Outside the pipeline. No upstream skill required. |
| propagated | /plate --hard invokes /hard-cheese <slug> after its final writing gate and before publication. |
At the verified-artifacts → share-for-review boundary. |
--hard propagates through /cheese → /mold → /cook → /press → /age → /cure → /plate. Upstream skills pass the flag; /plate is the only pipeline skill that invokes /hard-cheese.
Portability reference: ../cheese/references/harness-portability.md. It covers helper resolution, sub-agent dispatch, GitHub operations, and handoff transitions; prefer the bundled or repo-local helper first, and treat ${CLAUDE_SKILL_DIR} as optional host-provided fallback.
The handoff blocks below are the portable contract; slash commands are host renderings, not the control model.
-
Resolve scope.
diff_base = origin/main,diff_head = <short-sha of HEAD>.- If
.cheese/specs/<slug>.mdexists, load it as the intent reference (optional — diff is the ground truth). - Slug fallback when none supplied: the HEAD short SHA.
- If the working tree has no diff against
origin/main, exit0with"nothing to gate on"and write no artifact.
-
Freshness check. Check freshness before launching the gate:
python3 skills/hard-cheese/scripts/hard-cheese.pyz freshness-check \--slug <slug> --passing-score <n>Exit 0 (
previously_passed): print"previously passed"and exit0. Exit 2 (stale: HEAD moved or the last PASS score is below--passing-score) or 3 (new): continue to step 3. -
Compose the vibecheck prompt (faithful to Sankaranarayanan 2026, generalised to “share for review” so the gate stays implementation-agnostic):
Before this is shared for review, explain its causal logic in your own words. How does
work? Why does it produce the desired behavior? What state, control flow, or invariants does it rely on? Render a diff summary alongside the prompt. When invoked by
/plate, also render/plate’s final artifact inventory and{target, backend, verified}rows so the explanation covers the exact state about to be shared. -
Capture the user’s explanation as free text. No coaching, no example answers — the explanation is the artifact under test.
-
Spawn the judge sub-agent in fresh context (same pattern
/cook’s fan pathway uses for adversarial review). The judge:- Reads
references/judge-prompt.mdas its system prompt. - Receives the passing score threshold, the diff summary, the spec excerpt (if any), and the user’s explanation as context.
- Returns a JSON object:
{score, level, pass, feedback, socratic_qs}.
See
references/judge-prompt.mdfor the full system prompt and output shape.Skip this step when
--no-judgeis set: mark the attemptstatus: LOGGED, write the artifact, exit0. - Reads
-
On judge result:
score >= <passing-score>→ PASS.score < <passing-score>→ FAIL, render Socratic questions, loop to step 4 ifattempts < --socratic-cap. Judge error → ERROR attempt, print warning, exit0(fail-open — see## Divergence from the paper).
Append the attempt row:
python3 skills/hard-cheese/scripts/hard-cheese.pyz append-attempt \--slug <slug> --status <PASS|FAIL|ERROR> --score <n> \--feedback "<judge feedback>" --explanation "<user explanation>" -
On cap exhaustion: set the artifact
status: FAILED, print the path, exit non-zero. Downstream chains must not proceed.
Artifact
Section titled “Artifact”.cheese/hard-cheese/<slug>.md is the audit trail. The directory is gitignored by repo convention (.gitignore already ignores .cheese/), so the trail stays local — matching vibecheck’s local-only stance on telemetry.
Each file opens with a YAML frontmatter block that travels with the audit trail:
---slug: <slug>attribution: Sankaranarayanan 2026 / vibecheckrubric: SOLO Taxonomy (1-5), pass threshold = <passing-score>passing_score: <n>divergence: fail-open on judge error (vibecheck fails closed)diff_base: <sha>diff_head: <short-sha>status: PASS | FAIL | FAILED | LOGGEDattempts: <n>---The attempt log uses a 6-column markdown table (written by append-attempt):
| timestamp | head_sha | status | score | feedback | explanation || --- | --- | --- | --- | --- | --- || 2026-06-25T10:00:00+00:00 | a1b2c3d | FAIL | 2 | "Unistructural: lists steps but no causal link" | <user explanation verbatim> || 2026-06-25T10:05:00+00:00 | a1b2c3d | PASS | 4 | "Relational: explains why invariant holds" | <user explanation verbatim> |Attempts append; nothing is overwritten within a single invocation. If a re-invocation finds the artifact stale (HEAD moved), new attempt rows are appended below the prior ones — the trail is cumulative.
Sub-agent contract — fresh judge
Section titled “Sub-agent contract — fresh judge”- Fresh context, every invocation. Same-context judging is biased toward the code it helped write.
- Resolve a no-tool or read-only
reviewerthrough the shared agent resolver atdefaultpower andhigheffort. A general worker qualifies only with prompt-only no-write enforcement anddegraded: true. references/judge-prompt.mdis the system prompt. The judge reads the supplied diff summary, spec excerpt, and explanation, then returns JSON without repository writes.- JSON output is parsed. If parsing fails, the attempt is logged as
ERRORand the gate fails open (see## Divergence from the paper).
If the host harness has no sub-agent primitive, /hard-cheese is the wrong skill — the gate cannot run without a fresh judge. Recommend /hard-cheese --no-judge for users who still want the explanation captured as telemetry without the grading step.
Attribution
Section titled “Attribution”Sankaranarayanan, S. (2026). Mitigating ‘Epistemic Debt’ in Generative AI-Scaffolded Novice Programming using Metacognitive Scripts. Proceedings of the 13th ACM Conference on Learning at Scale. https://arxiv.org/abs/2602.20206
The implementation reference (intercept-at-acceptance, SOLO rubric, Socratic retry) is the open-source VS Code extension by the paper’s author:
https://github.com/sreecharansankaranarayanan/vibecheck
The attribution appears in this SKILL.md, in references/judge-prompt.md, and in every .cheese/hard-cheese/<slug>.md artifact so the citation travels with the audit trail.
Divergence from the paper
Section titled “Divergence from the paper”Hard-cheese departs from vibecheck in exactly one place, and the divergence is called out explicitly so it stays legible:
Vibecheck fails closed on judge error. If the Judge LLM cannot produce a verdict, the modal blocks code application until the judge recovers or the user retries with a different model.
Hard-cheese fails open on judge error. If the fresh-context judge sub-agent crashes, times out, or returns malformed JSON, the gate writes an ERROR attempt, prints a clear warning, and exits 0 — the user is allowed to proceed.
Rationale: judge invocation is per-PR-attempt and per-retry, and a strict fail-closed policy creates a worse experience under API hiccups than the epistemic-debt cost it averts. New divergences must be added here.
Composition with --auto
Section titled “Composition with --auto”--hard and --auto may coexist. The gate punctures auto exactly once inside terminal /plate --hard, after /plate verifies final artifacts and before publication. The user responds, then PASS permits publication, FAILED halts, and ERROR follows the documented fail-open behavior.
Commit-only /plate --hard does not fire because nothing is shared. For a new PR under auto, /plate honors explicit topology, infers an obviously cohesive single, and asks when stacked is recommended or shape is ambiguous. Non-TTY behavior lives in references/composition.md.
Output
Section titled “Output”When the gate ends, print:
Hard-cheese artifact: .cheese/hard-cheese/<slug>.mdStatus: PASS | FAILED | LOGGED | ERRORAttempts: <n>Followed by:
- On PASS:
Ready to share for review. - On FAILED:
Cap exhausted. Improve understanding of the change before sharing. - On LOGGED:
Telemetry only — judge skipped via --no-judge. - On ERROR: a one-line warning naming the failure mode and
Fail-open divergence active — gate exited 0; you may share for review at your discretion.
Preferred tools and fallbacks
Section titled “Preferred tools and fallbacks”| Need | Prefer | Fallback |
|---|---|---|
| Diff inspection for the user-facing summary | delta |
git diff --unified=3 |
| Reading the spec (when present) | bounded file read per code-intelligence-routing.md |
host file read |
| Spawning the judge | host sub-agent primitive (Agent() or harness equivalent) |
none — without sub-agent spawn, run --no-judge mode and tell the user the judge is unavailable |
| GitHub / PR context (out of scope here) | n/a | n/a |
- The judge sub-agent runs in fresh context. Do not let the same conversation that wrote the code grade the human’s understanding of it.
- Do not coach the user before they answer. The explanation is the artifact under test. Socratic questions appear only after a FAIL, and only the questions returned by the judge — no extra hints from the parent.
- Do not paraphrase the user’s explanation before passing it to the judge. The judge grades what the user wrote, verbatim.
- Do not skip the freshness check. Re-invoking after HEAD has moved must trigger a fresh attempt sequence — prior comprehension is stale once the code changes.
- Do not silently drop ERROR attempts. The fail-open divergence requires that every judge failure is recorded in the artifact and surfaced to the user as a warning.
- Do not invoke
/ghor any specific PR-creation tool. The gate’s contract is “before code is shared for review” — implementation-agnostic. - Apply the shared voice kernel (lives at
../age/references/voice.md): say what the gate result was, flag residual risk ascertain | speculating | don't know, do not soften FAILED into “almost passing”.
References
Section titled “References”references/judge-prompt.md— SOLO Taxonomy rubric, judge sub-agent system prompt, JSON output shape.references/composition.md— the full--hard/--automatrix and the single puncture point.skills/hard-cheese/scripts/hard-cheese.pyz freshness-check— checks whether a previous PASS is still fresh for the current HEAD and passing score (step 2).skills/hard-cheese/scripts/hard-cheese.pyz append-attempt— atomically appends an attempt row to the audit trail (step 6).
Agent resolution
Section titled “Agent resolution”Resolve the fresh judge through ../cheese/references/agent-resolution.md.
| Work | Preferred types | Permissions/isolation | Minimum power | Effort | Fallback |
|---|---|---|---|---|---|
| Grade the explanation | reviewer | no-tool or read-only, fresh-context | default | high | compatible reviewer, then general |
The canonical hard-cheese audit carries the shared agent_resolution block.
Composition — --hard and --auto
Section titled “Composition — --hard and --auto”--hard and --auto are both propagated flags. They may coexist. The gate fires at exactly one named point; everywhere else, each flag’s normal semantics apply.
Propagation graph
Section titled “Propagation graph”/cheese --hard → /mold → /cook → /press → /age → /cure → /plate --hard │ └──► /hard-cheeseEvery upstream skill is pass-through. /plate is the only pipeline skill that calls /hard-cheese, after it inventories, writes, and reads back all required durable artifacts.
The matrix
Section titled “The matrix”| Invocation | Gate fires? | When | Notes |
|---|---|---|---|
/hard-cheese <slug> standalone |
Yes | Immediately. | No pipeline state required. |
/plate --hard commit-only |
No | n/a | Nothing is shared for review. |
/plate --hard existing PR |
Yes | After final writes and validation, before update. | No layout question. |
/plate --hard new PR |
Yes | After topology resolution, final writes, and validation, before publish. | Explicit choices and cohesive singles skip the question; stack recommendations and ambiguous shapes ask under auto. |
Upstream --hard without terminal /plate |
No | n/a | The flag remains pending. |
The single puncture point
Section titled “The single puncture point”--hard punctures --auto exactly once inside terminal /plate, after the final artifact-writing gate and before publication. Intermediate cook, press, age, and cure phases do not pause.
Non-TTY guard
Section titled “Non-TTY guard”If /hard-cheese detects it is running without an interactive input stream (no human can respond to the vibecheck prompt), it fails closed and aborts. The puncture only makes sense when a human is in the loop. A vacuous “auto-pass” with no human present would defeat the entire mechanism.
Auto-driven CI pipelines should not pass --hard. If they do, the gate aborts with a clear error: "--hard requires an interactive TTY; remove --hard or run interactively".
Flag precedence summary
Section titled “Flag precedence summary”--autowithout--hard: chain runs forward;/plateapplies its new-PR review-shape policy.--hardwithout publication: no gate.--auto --hard:/plateresolves topology, verifies final writes, then fires the gate once.--auto --hardon a non-TTY: aborts with the documented error.
There is no silent precedence. The only point where one flag overrides the other is named here.
Judge sub-agent — system prompt and output shape
Section titled “Judge sub-agent — system prompt and output shape”This is the system prompt and contract for the fresh-context judge spawned by /hard-cheese. The parent skill loads this file, passes it as the sub-agent’s instructions, and parses the returned JSON.
Attribution
Section titled “Attribution”The rubric and threshold are taken from:
Sankaranarayanan, S. (2026). Mitigating ‘Epistemic Debt’ in Generative AI-Scaffolded Novice Programming using Metacognitive Scripts. Proceedings of the 13th ACM Conference on Learning at Scale. https://arxiv.org/abs/2602.20206
Implementation reference: https://github.com/sreecharansankaranarayanan/vibecheck
System prompt (verbatim — pass to the judge sub-agent)
Section titled “System prompt (verbatim — pass to the judge sub-agent)”You are a fresh-context judge evaluating whether a human author understands the causal logic of an AI-scaffolded code change they are about to share for review.
You have no prior context on this codebase, this author, or the conversation that produced the diff. That is intentional. Your job is to read the author’s explanation strictly on its own terms against the diff you are shown, and grade it against the SOLO Taxonomy of Observed Learning Outcomes (Biggs & Collis 1982), as adapted by Sankaranarayanan 2026 for AI-scaffolded code acceptance.
The SOLO levels (1–5):
- Prestructural — the response is irrelevant, restates the prompt, or misses the point entirely. The author has not engaged with the change.
- Unistructural — the response names a single element of the change (a file, a function, an output) without integrating it into a causal account.
- Multistructural — the response lists several elements of the change but treats them in isolation; no cause-and-effect linkage between them.
- Relational — the response explains how elements of the change interact: cause-and-effect is articulated, control flow and state are tied together, the author can defend why this change produces the desired behavior.
- Extended Abstract — the response generalises beyond the immediate change: invariants, trade-offs, what would change under different inputs, how this transfers to adjacent code.
Pass threshold:
score >= passing_score. Defaultpassing_scoreis 3 (Multistructural-or-higher).Per Sankaranarayanan 2026, the default threshold treats scores at or above Multistructural (3+ on this 1–5 scale) as sufficient causal understanding to defend the change in code review. The parent may supply a stricter or looser
passing_score; use that value for the booleanpassdecision while keeping the SOLO score and level faithful to the rubric. Scores belowpassing_scoreindicate the author has not yet met the configured gate. The Multistructural-vs-Relational distinction stays informative — a default level-3 pass with no cause-and-effect linkage is the minimum acceptable; a level-4 response is the aspirational target.Note on terminology: the paper labels the pass condition “Relational”. On this 1–5 mapping (Biggs & Collis), Relational is level 4 and Multistructural is level 3. The threshold rule above uses the level-3 label to stay unambiguous against the rubric; the paper’s “Relational pass condition” terminology and “score ≥ 3” are the same operational gate.
Grading rules — strictest reading wins:
- Steelman the strictest reading of the rubric. If the explanation is ambiguous between two adjacent levels, score the lower one. A generous judge defeats the gate’s purpose.
- Demand diff-grounded cause-and-effect. Template answers, generic restatements of “the code does X”, or descriptions that could apply to any code change are scored Multistructural at best. The explanation must cite specifics from the diff.
- Do not be charmed by fluent prose. Long, well-structured paragraphs that do not articulate causation are still Unistructural or Multistructural. Length is irrelevant; causal integration is everything.
- Do not infer understanding from absence. If the author omits a critical element (a control-flow branch, a non-obvious invariant), that omission lowers the score.
- The judge does not grade the code. The code may be wrong, weird, or suboptimal — that is
/age’s job. The judge grades the author’s understanding of the code as written.On FAIL (score < passing_score): return 2–4 Socratic questions that point the author toward the missing causal-logic component without revealing the answer. The questions should be specific to this diff and this explanation — not generic prompts. The goal is to provoke the author into the next attempt, not to teach them the code.
On PASS (score >= passing_score): return an empty
socratic_qsarray and a one-paragraphfeedbackfield explaining what the author got right.Output: a single JSON object, nothing else. No prose before or after.
Input shape passed to the judge
Section titled “Input shape passed to the judge”The parent skill sends the judge a single user message containing, in order:
- The configured
passing_scoreinteger (1..5; default3). - The spec excerpt (if
.cheese/specs/<slug>.mdexists) — up to ~30 lines. - The diff summary — files changed and key hunks, capped at ~80 lines.
- The author’s free-text explanation, delimited as a fenced block.
The judge does not request additional context. If the input is insufficient (no diff, no explanation), the judge returns score: 1, level: "Prestructural" with a feedback line explaining what was missing.
Output JSON shape
Section titled “Output JSON shape”{ "score": 1, "level": "Prestructural | Unistructural | Multistructural | Relational | Extended Abstract", "pass": false, "feedback": "one-paragraph critique grounded in the diff and the author's words", "socratic_qs": [ "specific question pointing at a missing causal-logic component", "second question, optional" ]}Constraints:
scoreis an integer 1–5.levelmatches the score exactly (1=Prestructural, 2=Unistructural, 3=Multistructural, 4=Relational, 5=Extended Abstract).passistrueiffscore >= passing_score.feedbackis a single paragraph, 2–5 sentences. No markdown headers, no lists.socratic_qsis an array of 2–4 strings on FAIL, an empty array on PASS. Each question ends with a question mark.
If the parent cannot parse the JSON, it treats the attempt as ERROR and applies the fail-open divergence — see skills/hard-cheese/SKILL.md ## Divergence from the paper.