SC

songjiun10-collab/senior-thinking-skills

开发工具
67 stars 质量 40 趋势 40

A skill collection that makes an AI think like a senior, principal, distinguished/fellow, and executive-level (CTO/VP-Eng) engineer before writing code — four stacked lenses (this task, cross-team,...

概览

A skill collection that makes an AI think like a senior, principal, distinguished/fellow, and executive-level (CTO/VP-Eng) engineer before writing code — four stacked lenses (this task, cross-team, company/industry-wide, organizational/business authority), most tasks only needing the first. Instead of one giant skill, it's split into . - — called directly by a person. Its job is orchestration. It can call the model-invoked skills below, but . - — a person can call it too, but the agent also pulls it out on its own when the task calls for it. The reusable disciplines live here. 1. — the router reads the situation and picks 2. — each one works independently. Want only the debugging discipline? Take just root-cause-discipline 3. — written assuming you'll reword it to match your team's conventions - beat one 1100-line skill. Only what's needed loads into context, and only the unneeded part can be deleted - Each skill's description carries .

README

Senior + Principal + Distinguished + Executive Engineer Mindset — Skill Bundle

A skill collection that makes an AI think like a senior, principal, distinguished/fellow, and executive-level (CTO/VP-Eng) engineer before writing code — four stacked lenses (this task, cross-team, company/industry-wide, organizational/business authority), most tasks only needing the first. Instead of one giant skill, it’s split into small, composable disciplines.

Structure

Split along the invocation axis:

  • user-invoked (router) — called directly by a person. Its job is orchestration. It can call the model-invoked skills below, but never calls another router.
  • model-invoked (discipline) — a person can call it too, but the agent also pulls it out on its own when the task calls for it. The reusable disciplines live here.

List

Router (user-invoked)

Skill Role
senior-engineer-mindset Picks the disciplines that fit the task and routes to them. Scale-adjustment and output format live here too

Disciplines (model-invoked)

Stage Skill Role
Verify search-first External APIs/libraries: docs, not memory
Verify context-economy Splits what belongs in context vs. in a file
Verify persistent-memory Captures a recurring correction/preference into a durable file (+ CLI) instead of re-learning it every session
Understand clarify-the-real-problem Digs out the real goal behind the request
Explore widen-the-solution-space Widens candidate approaches across 12 axes, narrows to 2-3
Decide weigh-tradeoffs Comparing alternatives + how much the decision matters (reversibility)
Decide record-the-why Records the decision and rejected alternatives permanently (ADR)
Design premortem Failure scenarios and failure detection
Design simplicity-budget YAGNI + complexity budget
Design design-for-the-next-reader Interfaces first + the reader six months out
Design interface-contracts Hyrum’s Law — anything you expose becomes a promise
Design threat-and-scale-check Trust boundaries · scale · layered defense
Verification verifiability-first Declaring success criteria + testable design
Plan bite-sized-plan File structure → task breakdown → 2-5 minute steps
Delegate delegate-to-subagents Judging when/how to delegate to a subagent, briefing, verifying (+ optional PreToolUse/PostToolUse hook bundle)
Debug root-cause-discipline Check if it’s already solved → root cause → evidence
Execution chestertons-fence Understand why something exists before removing or simplifying it
Execution surgical-change Don’t add lines that don’t trace to the request (+ change-size / severity labels)
Execution measure-before-optimizing Performance: evidence (measurement) first — never optimize on a guess
Execution honest-artifacts Unverified labels · reproducibility · metric traps
Look back fresh-context-review Strip the context you wrote it in, review on two axes: spec compliance · code quality
Look back verify-before-claiming Never claim done without evidence
Look back adversarial-review Interrogate with a disproof bias while course-correction is still cheap

How to use it

All three are valid:

  1. Install all of it and leave it alone — the router reads the situation and picks
  2. Install only the disciplines you need — each one works independently. Want only the debugging discipline? Take just root-cause-discipline
  3. Fork and edit it — written assuming you’ll reword it to match your team’s conventions

Design principles

  • 23 skills of 30-70 lines beat one 1100-line skill. Only what’s needed loads into context, and only the unneeded part can be deleted
  • Each skill’s description carries when to fire. The body carries what to do, nothing else
  • Scale-adjustment is the router’s job. Writing a design memo for a single function is a failure

Sources

This bundle pulled ideas from ten places and reconstructed them.

Source Link What was taken
addyosmani/agent-skills https://github.com/addyosmani/agent-skills Hyrum’s Law (API design), Chesterton’s Fence (understand before simplifying), disproof-biased review (adversarial framing), change-size thresholds (~100/300/1000 lines), severity labels (Critical/Nit/Optional/FYI), refusing “I’ll clean it up later”, measure→identify→fix→re-measure (performance optimization), ADR-style documentation that records the why
obra/superpowers https://github.com/obra/superpowers Most referenced — 4-phase systematic debugging + the 3-failure rule, a verification gate before “done”, spike/bounded/structural 3-track classification with a one-way ratchet, finely-sliced planning, red-flag/rationalization tables, YAGNI, complexity reduction, evidence-first — plus subagent delegation (recording a baseline before dispatch, reports/reviews as files, a capped fix→re-review loop with recorded rulings, the controller never self-fixes — from subagent-driven-development)
songjiun10-collab/Hncs https://github.com/songjiun10-collab/Hncs Unverified-value labeling, reproducibility audits, avoiding metric overfitting (favor the conservative choice), surgical changes, checking for an existing answer before digging deep, risk-tiered gates, subagent delegation principles (minimal briefing, one at a time, verify reports yourself — from CLAUDE.md’s “Controller/Implementer” section)
mattpocock/skills https://github.com/mattpocock/skills The whole structure — the user-invoked-router / model-invoked-discipline split, small composable skills, splitting spec-axis and standards-axis review
affaan-m/ECC https://github.com/affaan-m/ECC Checking docs before coding (research-first), fresh-context review, context economy (“optimize context, persist the rest”), the workflow loop
shinpr/sub-agents-skills https://github.com/shinpr/sub-agents-skills Splitting subagent permissions by role (review/implementation/testing need separated permissions, not just separated instructions)
WenyuChiou/agent-collab-skills https://github.com/WenyuChiou/agent-collab-skills Adversarial debate (putting two agents on opposing sides of a consequential decision), reconciling results from competing exploration
rohitg00/awesome-claude-code-toolkit https://github.com/rohitg00/awesome-claude-code-toolkit Multiple agents competitively exploring the same problem in isolated workspaces, then comparing and choosing
obra/superpowers (dispatching-parallel-agents) https://github.com/obra/superpowers/blob/main/skills/dispatching-parallel-agents/SKILL.md The bar for independent-domain parallel dispatch (several unrelated failures vs. related ones), parallel-brief structure (narrow scope + constraints + deliverable), pre-merge reconciliation steps (review summaries individually → check for overlap → run the full suite)
affaan-m/ECC (dmux-workflows) https://github.com/affaan-m/ECC/blob/main/skills/dmux-workflows/SKILL.md Uncommitted files are invisible under git-worktree isolation, and need to be added explicitly (seedPaths)

Per-skill provenance

Skill Primary source
senior-engineer-mindset (router) mattpocock (structure) + ECC (flow)
search-first ECC
context-economy ECC + Hncs
persistent-memory original — the explicit file-based workaround for the standing per-agent memory this bundle’s tooling doesn’t have, written after fact-checking xAI’s Grok Bot persistent-memory feature against official sources and confirming the gap is real
clarify-the-real-problem superpowers + Hncs
widen-the-solution-space original
weigh-tradeoffs original + Hncs (risk-tiered gates)
record-the-why addyosmani (documentation-and-adrs)
premortem original
simplicity-budget superpowers + Hncs
design-for-the-next-reader superpowers + mattpocock (deep modules)
interface-contracts addyosmani (Hyrum’s Law)
chestertons-fence addyosmani
adversarial-review addyosmani (doubt-driven-development)
threat-and-scale-check Hncs (layered defense)
verifiability-first superpowers (TDD) + Hncs (declaring success criteria)
bite-sized-plan superpowers + OpenAI Codex’s Plan Mode (validating precedent, not a source of new content — see Update log)
delegate-to-subagents superpowers (subagent-driven-development + dispatching-parallel-agents) + Hncs (CLAUDE.md “Controller/Implementer” section) + shinpr/sub-agents-skills (role-based permission separation) + WenyuChiou/agent-collab-skills · rohitg00/awesome-claude-code-toolkit (competing exploration, adversarial debate) + ECC (dmux-workflows, uncommitted-file handling under worktree isolation) + xAI’s Grok Bot product announcement (calibrating how often to interrupt for approval, named-agent status roster, reusable task templates that improve with corrections, coordinator/hierarchical delegation)
verify-before-claiming superpowers
root-cause-discipline superpowers (4-phase, 3-failure rule) + Hncs
surgical-change Hncs + addyosmani (change-size / severity labels)
measure-before-optimizing addyosmani (performance-optimization)
honest-artifacts Hncs
fresh-context-review ECC + mattpocock

What wasn’t brought in

Left out deliberately — skill docs alone can’t implement these, and including them would produce “documentation that pretends to work.”

  • ECC’s instincts / continuous-learning (auto-extracting patterns from a session), memory vaults, AgentShield — harness runtime features. persistent-memory (added later) isn’t a reversal of this call: it’s a manually-invoked script that actually writes and reads a real file, not a claim that the agent automatically remembers anything — the thing ruled out here was pretending an unbuilt automatic-memory feature exists.
  • Hncs’s hook-based enforcement (PreToolUse/PostToolUse gates) — needs a real execution environment
  • mattpocock’s issue-tracker-integration skills (triage, to-tickets) — outside this bundle’s scope (thinking before you write code)

Each original follows its own license. This bundle reconstructs those ideas, not a code copy.

Update log

  • 2026-08-24: Added measure-before-optimizing and record-the-why (both reconstructed from addyosmani/agent-skills — performance-optimization and documentation-and-adrs respectively). Re-searched GitHub and found two gaps in the original 20 (measuring before performance work, permanently recording decision reasoning). The rest of that repo’s skills (ci-cd, git-workflow, browser-testing, frontend-ui, etc.) stayed excluded — tooling/domain-specific, outside this bundle’s scope.
  • 2026-08-24 (later, same day): Re-swept the other 4 sources (superpowers, mattpocock, ECC, Hncs) — confirmed no gaps, nothing added. superpowers’ brainstorming (spike/bounded/structural 3-track classification + hard gate) is already the router’s own structure; mattpocock’s code-review (parallel spec-axis/standards-axis review) is already in fresh-context-review’s provenance; ECC’s security-review is specific to a Next.js/Supabase/Solana stack, out of scope and redundant with this session environment’s separate system skill. Widening the search to all of GitHub (including forrestchang/andrej-karpathy-skills, the most talked-about related repo) turned up the same result — its 4 principles turned out to already be the exact section titles under Hncs’s own CLAUDE.md “Working principles,” so nothing new there either.
  • 2026-08-24 (later, third pass): Added delegate-to-subagents (reconstructed from Hncs’s CLAUDE.md Controller/Implementer section — when to delegate, what to include/exclude in a brief, how to verify a report). The bundle’s first skill carrying scripts/ — check_dispatch_brief.py detects overlapping dispatches and transcript-paste-style briefs at Task-call time (PreToolUse/PostToolUse). Default is advisory (visible to a human, doesn’t affect Claude’s behavior); DELEGATE_HOOK_STRICT=1 actually blocks and returns the reason to Claude. Verified with 6 isolated-stdin cases (non-Task passthrough, clean brief passthrough, long-paste detection, count closing on PostToolUse, STRICT-mode exit 2, fail-open on malformed input) — explicitly noted in SKILL.md that this is a different design tier from Hncs’s real hooks (_hook_common.py, CRITICAL, deny-by-default).
  • 2026-08-25: Strengthened delegate-to-subagents — the first draft only looked at Hncs and missed the superpowers:subagent-driven-development skill that this very repo’s CLAUDE.md actually points to; read it and folded it in after that was flagged. Added: recording a baseline (git rev-parse HEAD) right before dispatch — use that baseline, not HEAD~1, when scoping a later diff/review (an intervening commit silently corrupts HEAD~1’s range); reports and reviews as files (survive compaction); a cap on the fix→re-review loop, with the controller adjudicating and recording a ruling once the cap is hit (never silently discarded); review findings go to one subagent at once rather than one subagent per finding; the controller never self-fixes; waiting for the completion signal instead of tight polling. superpowers’ full apparatus (git-worktree isolation, a ledger file, scripts/task-brief / scripts/review-package) didn’t fit this bundle’s 30-70-line compact-discipline style, so only the principles were kept, not the script infrastructure.
  • 2026-08-25 (later, same day): Widened the subagent-orchestration sweep per “check other projects too” — read the actual READMEs for shinpr/sub-agents-skills, WenyuChiou/agent-collab-skills, rohitg00/awesome-claude-code-toolkit. Folded in three things: (1) role-based permission separation (don’t give a reviewer write access — shinpr), (2) an explicit exception to “one at a time”: deliberately racing approaches in isolated workspaces for comparison (rohitg00 + WenyuChiou’s task-splitter/reconciler), (3) adversarial debate — putting two agents on opposing sides when judgment is split on a hard-to-reverse decision (WenyuChiou’s adversarial-debate). WenyuChiou’s shared-memory (.coord/memory.yml, a cross-session blackboard) overlapped heavily with the baseline-recording/ruling-recording principle already folded in from superpowers, so it wasn’t pulled in separately.
  • 2026-08-25 (later, third pass): Re-checked the remaining two sources on this same topic per “and ECC and matt-whatsit too.” Found nothing relevant in mattpocock (there was an X post about an /implement-spec skill implementing tickets in subagents at maximum concurrency, but it’s a tweet, not a maintained SKILL.md, so it wasn’t used as a source). Pulled one thing from ECC’s dmux-workflows (tmux-based parallel agent orchestration): under git-worktree isolation, uncommitted local files are invisible to a worker, and the seedPaths concept for adding them explicitly. In the process, discovered that superpowers itself has a dedicated skill, dispatching-parallel-agents, which had been missed — folded it in too: the bar for independent-domain parallel dispatch (dispatching several unrelated failures at once within one response) vs. sequential dispatch, and pre-merge reconciliation steps (review summaries individually → check for overlap → run the full suite). While integrating this, found that check_dispatch_brief.py’s “warn on any overlapping dispatch” logic was itself a problem — 2-3 concurrent independent-domain dispatches is exactly the normal pattern the skill recommends, and the hook was flagging all of it. Raised the threshold from “1 or more open” to “4 or more open” and softened the wording from “this is wrong” to “confirm these are actually independent.” Re-verified with isolated stdin (quiet through 3 open, warns on the 4th).
  • 2026-08-25 (later, fourth pass): Translated the entire bundle to English and blended Principal-engineer-level thinking into the existing Senior-engineer-level content, per explicit direction. Each discipline got at least one “Principal angle” bullet extending its existing checklist to cross-team blast radius, precedent-setting, and longer-horizon technical direction — not a rewrite into an org-strategy essay, just an added lens. Directory and name: frontmatter fields were left unchanged (English translation only touches prose and the description: field) so existing installs and cross-references keep working.
  • 2026-08-25 (later, fifth pass): Added a third tier above Principal — Distinguished/Fellow — per explicit direction (“a grade higher than Principal too”). Every one of the 22 disciplines got exactly one more bullet, always placed right after the existing Principal-level content, distinguishing scope by scale and time horizon: Principal reasons in terms of teams and quarters, Distinguished/Fellow reasons in terms of the whole company or industry and years (does this show up in company-wide strategy conversations, could it become a de facto standard or get open-sourced/talked about externally, does it need to survive scrutiny from someone joining in five years with zero institutional memory). Router retitled again to “Senior + Principal + Distinguished Engineer Mindset.” Same execution pattern as the translation pass — 5 parallel subagents on independent file batches, each verified after: grep confirmed exactly one “Distinguished” mention per discipline file (24 matches across 24 files — 22 disciplines + README + the router’s 3 intro mentions), all 23 skills re-validated with quick_validate.py, every name: field re-checked against its directory, no stray non-English text introduced.
  • 2026-08-25 (later, sixth pass): Added a fourth tier — Executive (CTO/VP-Eng) — above Distinguished/Fellow, from a brainstormed-options list the user picked “1, 2, 7” from. Distinguished/Fellow tops out as an individual-contributor lens (influence through technical credibility); Executive shifts to actual organizational levers: headcount and team structure, budget and total cost of ownership (not just engineering effort), and business risk (regulatory, competitive, customer-trust) a CFO or board would want visibility into. Every one of the 22 disciplines got one more bullet, placed right after the existing Distinguished/Fellow bullet, labeled “Executive angle (CTO/VP-Eng):”. Router retitled again to “Senior + Principal + Distinguished + Executive Engineer Mindset.” Same 5-parallel-subagent-batch pattern, same verification afterward: grep confirmed exactly one “Executive angle” per discipline file (22/22), all 23 skills re-validated.
  • 2026-08-25 (same brainstorm, in progress): The other two picked options are being worked in parallel — (1) a router-only trigger-description optimization pass via skill-creator’s scripts/run_loop (20 hand-written eval queries, 10 should-trigger/10 should-not, run against the live claude CLI in this environment since a browser isn’t available for the interactive eval-review step), and (2) splitting this bundle into its own standalone GitHub repository, pending the user’s answer on repo name/visibility/license.
  • 2026-08-25 (resolution): (1) The trigger-optimization pass ran 4 iterations and crashed on the 5th (claude -p subprocess error) - but the real finding came before the crash. Recall stayed at exactly 0% across 4 completely different description rewrites, which pointed at the harness rather than the description: skill-creator/scripts/run_eval.py registers the skill under test as a slash command file (.claude/commands/.md), not as a real model-invoked Skill - slash commands only fire on an explicit /name, so a natural-language query can never trigger one no matter how the description reads. None of the “improved” descriptions the loop proposed were applied, since they were optimized against an invalid measurement. (2) The bundle now also lives at its own repo: https://github.com/songjiun10-collab/Senior-thinking-skills (public, MIT-licensed, same 23 skills + README, split out of this Hncs checkout since it’s unrelated to the camera-science domain).
  • 2026-08-25 (later): Added two more principles to delegate-to-subagents, sourced from xAI’s public “Grok Bot” agent product announcement (not the unrelated reverse-engineered grok-bot-0.18-reconstructed GitHub repo initially misidentified as the source - that repo was explicitly declined as a reference since it redistributes a proprietary app’s original installers and reverse-engineered internals with no vendor authorization). The actual source was a product-announcement post describing two UX patterns: only interrupting for approval on consequential decisions rather than every step (“계속 보채지 않고, 필요한 것만 물어본다”), and a roster of named, persistent agents each showing a one-line status update rather than staying silent until done. Added as two new sections, “Calibrate how often you interrupt” and “Give long-running work a name and a status line.”
  • 2026-08-25 (later): Two more slides from the same product-announcement post added two more bullets to delegate-to-subagents: (1) “teach a task once, it remembers and repeats it, refining from corrections” → a bullet under “How to brief” about capturing a recurring brief as a reusable template that improves with corrections instead of being re-explained from scratch each time; (2) “a team-lead agent directs other agents, which coordinate with each other instead of the human approving every step” → a bullet under “Decide whether to delegate first” naming the flat-dispatch-vs-coordinator-subagent choice explicitly, reserved for jobs with real inter-agent handoffs rather than independent parallel work (which stays flat). The user separately pushed, across several turns, to also mine the unrelated b-nnett/grok-bot-0.18-reconstructed GitHub repo (an unauthorized reverse-engineered reconstruction of the same product, bundling the original app’s proprietary installers) for “just the ideas” behind its approval mechanism — declined throughout: paraphrasing insight derived from unauthorized reverse-engineering doesn’t launder its provenance, unlike this product-announcement post, which the vendor published themselves.
  • 2026-08-25 (later still): Elaborated the coordinator/team-lead pattern in delegate-to-subagents into its own “Coordinating a hierarchy” section — a two-level depth limit (no coordinator-of-coordinators), the coordinator needing its own narrow brief rather than an open-ended mandate, an explicit escalation path for a stuck worker, naming the hub-and-spoke-vs-shared-thread reporting-shape choice and defaulting to hub-and-spoke, applying verify-before-claiming one level up to the coordinator’s own reports, routing its status traffic to a file/log rather than live context, and a concrete signal for promoting a flat dispatch into this pattern retroactively (catching yourself hand-sequencing handoffs between agents).
  • 2026-08-25 (external review pass): Two rounds of external AI review against delegate-to-subagents and its bundled hook turned up real bugs, verified against the live Claude Code docs (code.claude.com/docs/en/tools-reference, .../hooks) and the script’s own behavior rather than taken on faith:
    • Wrong tool name. The hook’s matcher and its hardcoded tool_name check both said "Task" — current Claude Code names the subagent-dispatch tool Agent (confirmed against the live tools-reference page and this session’s own tool list); with the old matcher the hook silently never fired. Fixed in the settings.json snippet, the script’s tool_name check, and the SKILL.md prose, with a note for anyone on an older Claude Code version where Task was still correct.
    • State-file race condition. The old design was one shared JSON file, read-modified-written on every PreToolUse/PostToolUse call with no lock — concurrent dispatches (the exact case this skill tells you to run) could race and silently drop an entry. Redesigned as one marker file per open dispatch, named by the call’s tool_use_id (confirmed present and stable across Pre/Post in the live hooks docs); concurrent creates of distinct filenames don’t race the way shared read-modify-write did. Verified with an isolated 20-way-parallel stdin test — 20/20 markers survived.
    • PostToolUse closed dispatches FIFO, not by identity. The old code always popped the oldest open timestamp regardless of which dispatch actually finished, so completion order silently drifted from tracked state. Fixed by keying each marker file to its own tool_use_id, so PostToolUse removes exactly the dispatch that closed. Verified with an isolated test that opens 5 and closes the 3rd specifically, confirming the 3rd (not the 1st) is what’s gone afterward.
    • 1-hour staleness silently hid genuinely-long-running dispatches from the count, defeating the “too many concurrent” warning for exactly the long-running-agent case this skill’s own “name and a status line” section describes. Fixed by dropping the count-time staleness filter entirely — every open marker counts — and keeping only a much longer (24h) opportunistic-cleanup threshold for orphaned files from a crashed dispatch whose PostToolUse never ran.
    • Also fixed on inspection: extract_brief() now prefers the Agent tool’s actual prompt field over “longest string value” (which could misfire if another long string field existed), and the git-baseline / status-line / hook-scope prose was tightened to state its real limits explicitly — a git rev-parse HEAD baseline misses uncommitted changes already in the tree, “status line” needs an explicit file since dispatched agents have no live/streaming output channel, and the hook counts open dispatches, not file-level overlap (still your own job to check, per “never let two dispatches edit the same live file”).
    • Full isolated-stdin regression suite re-run after the redesign (non-Agent passthrough, clean brief, PostToolUse close, 4-then-5th-warns threshold, identity-correct close, long-paste detection, STRICT-mode exit 2, malformed-input fail-open, orphan cleanup, 20-way concurrency) — all passed.
  • 2026-08-25 (external review, round 2): Same reviewer re-checked the redesigned hook against the live file and confirmed the round-1 fixes hold (no more shared-state race, no more FIFO close, matcher/tool_name correctly Agent, prompt preferred over longest-string) — then found one new real bug in the redesign itself: a STRICT-mode block still created its marker file before exiting. open_dispatch() ran unconditionally, then finish(hit) exited 2 when hit and STRICT — but a blocked call never runs, so it never gets a PostToolUse to close the marker it just wrote. Every STRICT-blocked dispatch left a permanent phantom “open” entry (cleared only by the 24h orphan sweep), so repeated blocks could eventually trip the “too many concurrent dispatches” warning with zero agents actually running. Fixed by only creating the marker when the call will actually proceed (not (hit and STRICT)) — a plain warn-only hit (not STRICT) still creates its marker correctly, since that call does go on to run and will get a real PostToolUse. Verified with a new isolated regression pair: a STRICT-blocked call leaves 0 markers; the identical hit without STRICT leaves exactly 1, closed normally by its PostToolUse. Full suite re-run clean. Two smaller items from the same pass, both judged real but low-severity and left as documented limitations rather than “fixed” with more machinery: a dispatch missing tool_use_id (not expected on current Claude Code, per its hooks docs) gets warned about but silently untracked, and the 24h orphan-cleanup threshold can’t distinguish a crashed dispatch from a genuinely still-running one that old — both now stated explicitly in the script’s own docstring instead of glossed over.
  • 2026-08-25 (Grok Bot fact-check pass): The user asked for actual research on Grok Bot rather than another pasted screenshot. Searched and confirmed via independent outlets (Unite.AI, VentureBeat, Composio, Reworked — xAI’s own official launch, Aug 11 2026, distinct from the unrelated reverse-engineered grok-bot-0.18-reconstructed repo that stays off-limits per earlier entries) that Grok Bot does have genuine persistent per-bot memory (“the more you correct and guide a bot, the better it fits how you work”) and a watch-once-then-replay-on-a-schedule workflow-capture feature — confirming the read from the prior round that these are real product features, not just marketing gloss, and that they still don’t map onto Claude Code’s Agent tool (each dispatch starts with zero session history, confirmed again against nothing in this research contradicting it). One piece turned out to generalize honestly: Grok Bot’s “replay on a schedule” matches a capability Claude Code environments can actually have (a cron-style trigger that re-fires a saved prompt) — added one sentence to the reusable-template bullet under “How to brief” connecting the two, explicit that it only applies where such a scheduler exists. Nothing else from this pass was added — the rest (persistent bot memory, multi-bot group chat, granular per-action approval, always-on cloud execution) was either already covered by existing sections or, per the prior round’s finding, not something Claude Code’s dispatch mechanism actually has.
  • 2026-08-25 (new skill: persistent-memory): The Grok-Bot fact-check pass left one thread open — the honest workaround for standing per-agent memory (“write it to a file, load the file”) existed only as one bullet inside delegate-to-subagents, without an actual mechanism. Built it out into its own 23rd skill: when to make a memory file vs. put it in CLAUDE.md vs. skip it entirely, what belongs in it (distilled corrections, not transcript), and how to keep it honest (review periodically, resolve conflicting entries, delete what’s stale — same honest-artifacts logic applied to the memory file itself). Bundles scripts/memory.py, a dependency-free CLI (show/append/list over one markdown file per topic under .claude/memory/) — verified with 8 isolated tests (empty list, missing-topic show, two appends, topic-name sanitization, empty-note rejection, no-args/bad-arity usage+exit-1). Explicitly not a reversal of the earlier “What wasn’t brought in” call against ECC’s auto-extracting memory/memory vaults — that call ruled out claiming an unbuilt automatic-memory feature exists; this is a manually-invoked script that actually reads and writes a real file, the same category as the bundle’s other scripted skill (delegate-to-subagents’s hook). Cross-referenced from delegate-to-subagents’s reusable-template bullet and added to the router’s skill table and situational-picks table. All three touched skills (persistent-memory, delegate-to-subagents, senior-engineer-mindset) re-validated with quick_validate.py; skill count in this README updated 22 → 23.
  • 2026-08-25 (Codex Plan Mode cross-reference): User asked about a specific Codex feature - breaking a complex task into parts and showing that breakdown. Verified directly against OpenAI’s own developer blog (developers.openai.com/blog/run-long-horizon-tasks-with-codex) rather than answering from memory: Codex has a real “Plan Mode” (/plan), which breaks a task into a reviewable step sequence with acceptance criteria before touching code, asking follow-up questions first when something’s unclear - available in the Codex app, CLI, and IDE extension. This validates bite-sized-plan’s existing approach rather than adding new content to it (this bundle already does the same thing, one level more granular - 2-5 minute steps, not just phases) - added one sentence citing the precedent, not a new section. quick_validate.py re-run clean.
  • 2026-08-25 (new bundled hook: senior-engineer-mindset gets one too): User asked whether an AskUserQuestion check-in could be mechanically nudged for brainstorming-adjacent, hard-to-reverse decisions, choosing advisory-only over a hard block when given the tradeoff explicitly (false-positive risk from a path-based heuristic, plus consistency with delegate-to-subagents’s own hook, which is also advisory-by-default). Added scripts/check_ask_before_hard_change.py, the router’s first bundled hook: a PostToolUse on AskUserQuestion marks a fresh timestamp sentinel; a PreToolUse on Edit|Write|MultiEdit warns (doesn’t block, unless ASK_HOOK_STRICT=1) when the target path looks like a hard-to-reverse surface (schema/migration/public-API-spec/dependency-manifest - the same category the router’s own “Situational picks” table already names) and no check-in happened in the last ASK_HOOK_WINDOW_SECONDS (default 1800s). Verified with 9 isolated stdin tests (unrelated file silent, schema file warns, AskUserQuestion marks fresh, same file now silent, migrations/ path, sentinel expiry re-triggers the warning, STRICT mode exit 2, malformed-input fail-open, dependency-manifest match). Documented as a path-based heuristic, not real reversibility judgment, in both the script’s own docstring and the SKILL.md section - same honesty framing as delegate-to-subagents’s hook about its own limits.
  • 2026-08-25 (delegate-to-subagents: communication structure beyond one-shot dispatch): User asked how to design the communication structure with subagents. Rather than answer generically, checked this session’s own actual tools first (SendMessage/ListAgents, loaded via ToolSearch and read in full) instead of guessing - the whole skill up to this point assumed strictly one-shot dispatch (brief in, wait, read result, no further contact). Added “Talking to an agent that’s already running”: continuing a named agent instead of re-dispatching it from scratch (resumes from its own transcript with full context), subscribing to a completion notice instead of polling (the concrete mechanism behind the existing “wait for the completion signal” advice), naming the coordinator’s “shared thread” option as literally this capability, and a new safety point not previously in the skill - never ask a peer agent/session to do something your own session was denied or blocked from doing (cross-session permission laundering). Explicitly scoped as an addendum, not a replacement - noted that plain single-session Claude Code doesn’t have this capability set. Generic (not Hncs-specific), so applied to both the portable bundle and manually re-applied to the Hncs fork.
  • 2026-08-25 (new script: execution_manager.py, real code infrastructure): User’s diagram-based request for a “harness” turned out, after AskUserQuestion narrowed scope, to mean actual code (not more skill prose) - explicitly scoped down from “auto-detect problems and auto-redirect workers” (not really buildable; problem-detection and redirection judgment are inherently LLM-judgment steps, not something a script can do) to “status tracking + staleness warning,” with the judgment call left to whoever’s coordinating. Added scripts/execution_manager.py: a CLI (start/update/dashboard/clear) tracking each worker’s status (in_progress/blocked/complete) in one JSON file, plus an optional hook mode (wired to PreToolUse/Agent, same matcher block as check_dispatch_brief.py) that surfaces a warning only when something’s been blocked past EXECUTION_MANAGER_BLOCKED_WARN_SECONDS (default 1800s) - not a gate, purely advisory like this bundle’s other scripts. Verified with 13 isolated CLI/stdin cases (empty dashboard, start/update/clear, unknown-status rejection, hook silent when nothing’s stale vs. warns once staleness is simulated, PostToolUse and non-Agent calls both silent, malformed stdin fails open, mixed-state multi-worker dashboard). Documented in a new “Optional: track execution state with a script” section, and cross-referenced from “Coordinating a hierarchy”'s existing “route status traffic to a file” bullet. Generic, applied to both the portable bundle and manually re-applied to the Hncs fork.
  • 2026-08-25 (new section: delegate-to-subagents “Scheduled/unattended execution” — verified live, not from a claim): User asked to actually imitate Grok Bot’s “always-on cloud computer” angle, narrowed via AskUserQuestion to extending this session’s real scheduling tools (create_trigger/send_later) rather than more Grok-Bot-attributed prose. Ran a real demo instead of writing the section from assumption: scheduled a one-shot trigger 2 minutes out (send_later, which wraps create_trigger’s self-bind run_once_at), waited for the actual completion notification (no polling), then ran the scheduled command for real - a persistent-memory append that landed in .claude/memory/scheduled-execution-demo.md, confirmed by reading the file back. Documented the section citing this real run, not a hypothetical, and explicitly distinguished “unattended re-invocation at a point in time” from Grok Bot’s stronger claimed “always running” behavior - the container this fires into can still be reclaimed between firings. Cross-referenced from the existing “reusable template + scheduler” bullet under “How to brief”. Generic, applied to both the portable bundle and manually re-applied to the Hncs fork.
  • 2026-08-26 (Grok Bot official docs pass — docs.x.ai, not a screenshot or third-party summary): User pointed directly at xAI’s official Grok Bot documentation (docs.x.ai/grok-bot/* — overview, get-started, use-cases, skills-routines-and-automations, computer-and-apps, approvals-security-and-privacy). This confirmed the earlier suspicion held throughout this bundle: none of these official pages mention the internal-implementation-level field names (sourceEvidenceIds, parentAgentToolCallId, etc.) from the original unattributed GPT-researched list — that content stays declined. Two genuinely new, generic, independently-justifiable principles were pulled from the official pages and added to delegate-to-subagents: (1) under “Scheduled/unattended execution” — a routine should get one manual test-run (input actually picked up, output lands where expected, failure visible, not swallowed) before being trusted unattended, matching the official docs’ own “run test executions to verify input selection, output format, audit trails, approval checkpoints, and explicit failure states” requirement before activating a routine; (2) under “Calibrate how often you interrupt” — a confirmation gate only stops the next action, it doesn’t undo one that already ran, matching the official docs’ “An approval controls the proposed action. It does not reverse work already completed.” A third candidate (skills created by “teach once via a recorded browser demonstration”) was declined as non-portable — it depends on Grok Bot’s computer-use screen-recording feature, which has no equivalent in this environment. Generic, applied to both the portable bundle and manually re-applied to the Hncs fork.
View this README on GitHub

推荐工具

换一个关键词,或者移除筛选条件。

安装

npx skillfish add songjiun10-collab/senior-thinking-skills