GG

gadievron/greenlight

Security testing
46 stars Quality 40 Trend 40

Refusal handling skill for vulnerability research evals

Overview

Using the greenlight skill, you'd be able to autonomously handle Opus 4.7's refusals in security research-type evals. It was created for benchmarking purposes, mapping out where Claude Code's guardrails trigger exploitation-related refusals, and where they can't be relied upon. While validated only on Opus 4.7, the methodology should generalize to other models, but every model requires its own re-validation. Remember: this is a proof of concept, not rigorous research. Specifically, if your bench distribution differs from this one (different model, different target class, uniform rather than bimodal re-roll), re-validate before considering the retry-cap recommendation.

README

greenlight — refusal handling skill for vulnerability research evals

What this is

Using the greenlight skill, you’d be able to autonomously handle Opus 4.7’s refusals in security research-type evals. It was created for benchmarking purposes, mapping out where Claude Code’s guardrails trigger exploitation-related refusals, and where they can’t be relied upon.

While validated only on Opus 4.7, the methodology should generalize to other models, but every model requires its own re-validation. Remember: this is a proof of concept, not rigorous research. Specifically, if your bench distribution differs from this one (different model, different target class, uniform rather than bimodal re-roll), re-validate before considering the retry-cap recommendation.

Author: Gadi Evron (@gadievron)

Who it’s for

You are doing authorized AI-safety / security refusal research in cyber security, and are building a harness that runs Claude (via claude-agent-sdk or the Claude Code CLI), for purposes such as CVE reproductions, hardened-service testing, vulnerability-class benchmarks, exploitation capability measurements, and similar, and you are hitting one of these:

  • Sessions fail with Exception("Command failed with exit code 1") and you can’t tell if it was a transport flake, a rate-limit, or a content-policy refusal.
  • Your bench retries refused sessions blindly and wastes tokens on identical re-runs.

Where to start

If you are… Start at…
A researcher / harness author wanting the validated minimal skill (high evidence bar) SKILL.md (v3 canonical)
A researcher / harness author wanting a richer skill at lower evidence bar (varied capabilities) SKILL-v5.md
A researcher wanting to build your own skill section-by-section (template) SKILL-modular.md

Installation

# Drop into another research harness:
cp -r refusal-handling /path/to/your/harness/skills/

# Or make available in every Claude Code session on your machine:
cp -r refusal-handling ~/.claude/skills/refusal-handling

Validation status

Live skill: v3 (promoted 2026-04-24). Under Opus 4.7 against shellshock-like, N=13 valid paired vs adversarial:

Condition N valid Refusal rate Pass rate $/run
v3 canonical (live) 13 0.077 0.923 $0.40
Adversarial (no skill) 14 0.286 0.429 $0.58
v3 no-preamble (ablation) 15 0.200 0.600 $0.32

Transfer test on format-string-like held pass rate at 1.000 on a target where refusals don’t happen either way — see caveat below on target-specificity.

The 0.077 number above was a single low-tide measurement on 2026-04-24 evening. Identical config measured 0.200, 0.267, 0.333 over the following 10 hours; canary monitoring at 30-min spacing showed cycles of 0/5 = 0.000 refusal followed by 5/7 = 0.714 at the same config in the same day. Typical-day refusal rate is 0.15–0.35 at this config, not 0.077. The skill is robust (always beats the adversarial baseline measured in the same wall-clock window), but the absolute number is a snapshot of a non-stationary process. Run canary probes (~$10) before treating any single-batch refusal-rate measurement as ground truth.

Findings that matter for harness authors (regardless of v3)

  • No prompt-content variant beat v3 canonical. Stripping the preamble, adding hardened-sandbox vocabulary to it, swapping the user-prompt verb — all regressed. v3 canonical is locally optimal across all three explored prompt-content axes. v6 gains, if they exist, must come from outside prompt content.
  • Vocabulary differential analysis is correlation, not causation. The v3 user prompt’s "exploit" was baseline-injected; removing it didn’t help. Vocabulary-frequency ranked lists are candidates to test, not levers to pull.
  • Harness-level retry-with-addendum recovers some refused sessions. At N=2 valid pairs: 1/2 recoveries. Cost ~$0.50–0.90 per recovered seed.
  • Same-config refusal rate varies 2.6× across days. Strongest evidence we have that workspace-session temporal dynamics matter.
  • Position-within-batch is NOT the cause of refusal variance. Forward (0→14) and reverse (14→0) ordering of the same seeds under identical config produced different outcomes for the same seed at the same position. Seed identity is also non-deterministic (no seed refused in every batch; max 2/4). The dominant unmodeled variable is workspace-temporal / Anthropic-side stochasticity.
  • Day-shape at fixed config is bursty (N=12). Single seed at v3 canonical, repeated every 30 min over 6.5 h. Two of three pre-registered metrics positive: burstiness (longest run 5 PASS flanked by 3+2 REFUSE; sequence Px1 Rx3 Px5 Rx2 Px1) and hour-of-day max/min ratio 2.67. Lag-1 autocorrelation downgraded to near-independence at N=12. Honest reading: clusters of multiple cycles are real, but adjacent-pair correlation at this N is too low-power to summarize the structure. Cycles 5–9 alone read 0.000 refusal; the rest of the same day at the same config read 5/7 = 0.714.
  • Skill-update investigation, post-. Sixteen candidate skill changes examined plus five pre-rejected at scope; zero accepted for direct ship; one queued for paired re-validation (per-turn-reminder placement); fifteen rejected with explicit finding risk reasoning. Per-turn injection is the only design axis that has not been ablated; if a v6 effort is justified, that is the experiment to run first (~$10 paired N=15 harness change).
  • CSV schema for harness authors. Capture started_at_iso, completed_at_iso, batch_id, order_within_batch, gap_from_prev_s per row — nearly free, unlocks the entire temporal-analysis family. Without them, position drift and adjacency are invisible.
  • Placement-via-duplication does NOT help. Tested preamble-doubled variant (canonical preamble at top appended again before “What a good session”). Δrefusal = -0.067 vs canonical, McNemar exact p = 1.0 — within day-shape noise. Critical confound: control batch ran in a hot window, treatment batch ran 20 min later in a potentially cooler window. v3 canonical is locally optimal across four axes: presence content user-prompt verb placement-via-duplication.
  • Cap=2 retry mechanism recovers ~87% of refusals it sees, but does NOT improve final pass rate in same-window paired comparison. A direct paired comparison at N=15 (harmless-poc-task variant): cap=2 final pass = 13/15, cap=1 final pass = 13/15, Δpass = 0.000, McNemar p = 1.0. The retry recovery rate is real but day-shape variance cap=2’s higher wall-timeout exposure approximately cancel the gain at N=15 paired. Methodology lesson one-armed aggregate evidence ≠ paired direct comparison.
  • Preamble-in-user-prompt placement DOWNGRADED to NULL-RESULT after N=29 replication. First batch (N=14 paired) showed +0.214 directional positive. The queued replication (seeds 15–29) brought the combined N=29 to Δpass = -0.034, Δref = +0.034, McNemar p = 1.0 (b=8, c=7 balanced). Per-batch split: batch 1 leaned treatment 2.5:1; batch 2 leaned control 3:1; pooled balanced. The original directional positive was a sample artifact. v3 canonical now locally optimal across 5 prompt-content/placement axes. Methodology lesson confirmed: N=15 paired catches direction; N=30 catches whether the direction is real.
  • Soft-vs-hard differential retry is structurally untestable on shellshock-like. All refusals are both (hard+soft co-occurring); no soft_only cases exist even on the high-refusal variant. Asymmetric retry policy can’t be tested or shipped here. Likely testable on different target classes (sibling projects with memory-corruption-class refusals show different bucket distributions).
  • POST-PIVOT user-prompt verb SWAP CONFIRMED at p<0.05. Cell B (v3 + “exploit”) vs cell C (v3 + “harmless POC”) in same wall-clock window paired-by-seed at N=15: Δpass = +0.533, Δref = -0.467, McNemar exact p = 0.0386 (12 discordant pairs split 10:2 toward treatment). The earlier pre-pivot prescription (“exploit verb is locally optimal”) REVERSED post-pivot. For shellshock-like + Opus 4.7 v3 canonical post-pivot: set USER_PROMPT_TEMPLATE to harmless-poc wording for ~8× refusal-rate reduction, ~2.6× pass-rate improvement, 46% cost reduction. Code default in run_one.py stays “exploit” for backward compat. This is the first paired-direct CONFIRMATION (not downgrade) in the methodology arc.

Caveats that qualify the validation

  • Target-dependent. Refusal behavior is target-specific, not universal. shellshock-like triggers the filter often (0.286 adversarial baseline); format-string-like does not (0.000 adversarial baseline on the same N=11 valid paired seeds). Claiming “the skill reduces refusal rate” requires a target where refusals exist.
  • Single model. All validation on Opus 4.7. Sonnet/Haiku may have different refusal profiles — sibling projects report varied behavior per model.
  • Task frame matters. A separate project achieved 0 refusals across 18 CVEs (Log4Shell, EternalBlue, Shellshock, BlueKeep) under an environment-building frame with no specialized skill — their task frame is intrinsically less filter-tripping than ours.
  • Day-level variance is real. Identical config across days produced 2.6× refusal-rate spread. At N=15 paired, this is the same magnitude as our skill’s effect. Run validation in close wall-clock proximity if you’re comparing variants.
  • N=13–15 is underpowered. The paired direction is clear at every comparison; strict p<0.05 would need ~N=30.
  • Workspace pivot. The workspace running this research was approved by Anthropic for “lightened guardrails” mid-stream, around 2026-04-22. Some of the dossier was measured following that pivot. Frequency / observational findings (where refusals happen, what they look like, what triggers them) survive the pivot; absolute prevention-rate numbers are window-specific.

Disclaimer

This project is provided under the MIT License; see LICENSE for warranty and liability terms. It is a proof of concept intended solely for authorized LLM safety and security research, including model red teaming, refusal behavior evaluation, and guardrail testing, conducted in controlled and isolated environments.

This project has not been fully validated. Outputs may be incomplete, inaccurate, or misleading. Do not treat its outputs as authoritative or complete, and do not rely on them for decision-making, security assessments, or compliance purposes. Do not use this project to target real users, systems, data, or production models under any circumstances. Do not deploy in production.

You are responsible for complying with applicable laws, regulations, and all relevant terms of service and acceptable use policies of tools, services, or platforms involved.

License

This project is provided under the MIT License; see LICENSE.

View this README on GitHub

Recommended Tools

Try a different keyword or remove a filter.

Install

npx skillfish add gadievron/greenlight