Loading instructions when they matter

N17Q loaded specialized instructions, tools, references, and checks only when a task earned them, preserving context and authority boundaries around reusable expertise.

I put every N17Q operating instruction into one agent prompt.

Repository work, source review, document export, replay, MCP integration, approval, sandbox recovery, and trace evaluation all appeared before a simple file task. The model had access to the right knowledge somewhere inside a long undifferentiated manual.

It chose a replay instruction during a live-like record task and spent attention on tools that were not available.

Reusable expertise needed progressive disclosure. I began treating each specialized workflow as a skill loaded only when the task and policy justified it.

A skill was more than prompt text

In N17Q, a skill package could include scoped instructions, terminology, decision steps, reference documents, schemas, templates, deterministic checks, and requests for particular product capabilities.

It did not carry credentials or unconditional tool authority. The run compiler decided which capabilities were eligible independently.

The package had identity, version, owner, supported task types, data assumptions, required inputs, output contract, and test fixtures.

Expertise became a versioned component rather than copied prose.

The task classifier proposed one or more relevant skills from goal, target, workspace state, and user request. Policy filtered by workspace, data class, trust, and environment.

The model could request another skill after discovering a need. It could not install or activate one automatically.

N17Q recorded candidate, selected, loaded, and used as distinct states. A skill absent from context also could not supply tools through a hidden route.

Progressive disclosure began with an observable decision.

The base agent stayed small

Core context covered task, authority model, current workspace, tool-use protocol, safety state, budgets, and how to ask for missing expertise.

Specialized details lived outside until needed. This made the base prompt easier to review and reduced conflict between unrelated workflows.

The base still enforced universal principles such as proposal before consequence, current policy, evidence, and stop conditions. Skills could refine and never weaken them.

A small foundation made specialization easier to trust.

A repository skill could apply only inside one workspace and one task phase. A citation-review skill could govern evidence assessment without authorizing document mutation. An export skill could prepare an artifact and not deliver it.

N17Q compiled instructions with scope and precedence. When two skills overlapped, conflicts became explicit or one higher-level workflow orchestrated them.

Text from a skill did not outrank product policy because it sounded specialized.

Expertise had boundaries as well as content.

Tools were requested, not smuggled

A skill declared capabilities it could use and why. The run compiler intersected that request with registry, task, requester authority, environment, data policy, and budgets.

The resulting tool set entered the context manifest. A missing tool produced a supported degraded path, a capability request, or a stop.

The skill could not include an executable binary or server and thereby bypass integration review. Code assets ran only through reviewed sandbox capabilities.

Instruction packaging never became a covert permission system.

Some skills included a short procedure and a larger reference library. Sending every reference defeated the purpose.

N17Q loaded the main instructions fully, then selected linked references by task subproblem and explicit routing. The manifest named what entered context and what remained available.

Required legal, security, or migration caveats could be pinned. Optional examples stayed on demand.

Progressive disclosure happened within the skill as well as between skills.

A skill could be wrong for the version

Frameworks, APIs, provider tools, and project conventions changed. A reusable workflow without version constraints could introduce anachronistic or invalid instructions.

Skill metadata named compatible environment, tool contracts, repository patterns, and reference revisions. Selection failed or requested migration when the workspace did not match.

Historical replay pinned the skill version originally loaded. Counterfactual runs could change it explicitly.

Versioned expertise made advice reproducible and replaceable.

A reviewed skill was more trusted than arbitrary repository text and still not executable policy. It could contain a mistake, ambiguous example, or compromised reference.

Tool requests passed normalization and current policy. Outputs passed validation. World-state evaluation checked consequences.

The registry recorded provenance, review status, content digest, and last fixture run. High-consequence workflows required stronger review.

Trust informed selection without eliminating containment.

Skills reduced tool-description overload

The monolithic agent saw every capability description on every turn. A selected skill could request a small task-shaped set and provide the vocabulary needed to use it correctly.

The repository-review skill saw file, search, patch, and test capabilities. The effect-recovery skill saw receipts and bounded status query. Neither needed the other's full catalogue.

This saved context and reduced alias confusion.

Expertise and affordances arrived together under policy.

A structured request named unresolved goal, missing knowledge, candidate skill, expected inputs, and why current methods were insufficient.

N17Q showed the request to the owner or applied preapproved selection policy. Loading still created a context transition and consumed budget.

Repeated requests for an ineligible skill shared denial lineage. The model could ask for manual guidance instead.

Self-awareness of a gap did not become self-installation.

Skill composition needed a plan

Loading source research, repository migration, and release review at once could recreate the monolith in pieces.

N17Q preferred phases. One skill produced an evidence artifact. Another consumed it to prepare a patch. A review skill inspected the resulting world and receipts.

The checkpoint connected outputs under stable identity. Each phase received the minimum active skills and tools.

Composition followed workflow state rather than accumulating permanent context.

Two skills could prescribe different command order, file format, evidence standard, or review policy. N17Q detected overlapping scoped directives and their precedence metadata.

It did not ask the model to silently blend security-critical conflicts. The compiler selected one under explicit project rule or paused for owner decision.

The context manifest preserved the resolution. Updating a skill could re-open compatibility review.

Reusable expertise needed conflict management to remain reusable.

Outputs had contracts

A skill did not simply promise to help. It declared what artifact or decision state it produced: evidence set, normalized patch, scenario fixture, approval object, trace assessment, or handoff note.

Deterministic validators checked shape and product semantics. The model could produce prose around the artifact and could not mark it accepted.

Downstream skills consumed verified identities rather than scraping the previous response.

Structured handoff kept specialization from becoming a chain of lossy conversations.

Skill tests included representative tasks, missing inputs, wrong environment versions, hostile repository text, denied tools, budget exhaustion, and ambiguous outcomes.

They asserted world state, tool scope, artifact contract, policy compliance, and final account. A model reciting the procedure without producing a safe result did not pass.

Counterfactual runs compared skill versions against the same fixture world. Missing fixtures failed closed.

Expertise earned confidence through behavior under constraints.

Skill loading entered the trace

N17Q recorded selection reason, package and reference digests, instructions compiled, capabilities requested and granted, context cost, output artifacts, and unload state.

This let evaluation distinguish a model failure from missing or conflicting guidance. It also exposed whether a supposedly unused skill had influenced tool availability.

The final account could name specialized methods without pretending they were people.

Progressive disclosure remained observable.

Once a phase ended, its detailed instructions and tools did not need to remain in every context. N17Q removed them at a checkpoint while preserving produced artifacts and history.

An unresolved effect kept its recovery contract active even if the originating skill unloaded. Authority and safe cleanup outlived instructional convenience.

Reloading later used a current or pinned compatible version and explicit transition.

Context could shrink without forgetting consequence.

Skill updates were migrations

Changing decision steps, examples, schemas, or required tools could alter behavior. The package received a new version and ran compatibility fixtures.

In-progress runs stayed pinned or migrated at a checkpoint with visible differences. Old summaries did not silently receive new meaning.

Revoking a compromised skill prevented new loading and stopped pending tasks that depended on it. Historical artifacts remained for audit under retention.

Lifecycle completed the architecture around reuse.

N17Q could list installed or available skill packages by name, purpose, version, provenance, supported tasks, and required capabilities. Discovery made them candidates.

Activation required an approved package identity and current compatibility. A same-named replacement did not inherit trust. Unknown remote packages never entered model context automatically.

The registry snapshot joined the run so later evaluation knew what expertise could have been selected.

Skill discovery followed the same boundary as MCP tools.

Names and descriptions were not routing authority

A package could call itself “safe release expert” and claim universal applicability. N17Q used reviewed metadata, task predicates, compatibility, and policy to select it.

Names remained useful labels. Description changes received version review because they influenced classification and model requests. A persuasive paragraph could not broaden tool or data scope.

Routing evidence and execution authority stayed outside self-description.

Reusable knowledge did not certify itself.

Acceptance rate or one benchmark score could reward a skill that produced conventional answers. N17Q tracked scenario completion, hard invariants, tool scope, evidence, artifact quality, budget, and account accuracy by version.

Human review could assess clarity and usefulness with cited artifacts. A strong result in repository migration did not imply fitness for security review.

The selector used declared compatibility and explicit policy before quality preferences.

No global “best skill” number decided every task.

Fallback was a product choice

If a skill was missing, incompatible, or failed to load, N17Q could continue with core instructions, choose a reviewed alternative, ask the owner, or stop.

It never substituted another package with broader data needs or tools silently. The context transition explained the fallback and its limitations.

The artifact contract remained the same where compatibility was claimed. Otherwise the task phase changed explicitly.

Graceful degradation preserved authority and expectations together.

An owner could supply a one-off checklist or correction for the current task. N17Q stored it as scoped human instruction, not a permanent skill automatically.

Promoting repeated guidance into a package required review, metadata, references, tests, and version. Private notes did not become model-training or shared registry material by default.

This kept local expertise useful without turning every intervention into invisible global behavior.

Reuse was deliberate and attributable.

A stable procedure could depend on changing framework documentation or policy. Skill metadata attached freshness and authority requirements to references.

Loading checked current eligible revisions. Historical replay pinned old sources. A stale reference could disable only the affected path or request a current retrieval capability.

The package did not become fully obsolete because one link changed, and it could not cite an old API as current silently.

Procedure identity and source currency remained distinct.

A repository skill might need an accessibility review after rendering. It could propose the next phase and output the required artifact.

It did not recursively load another package inside hidden model context. N17Q selected the next skill, capabilities, budget, and checkpoint under current policy.

The trace showed the transition and prevented nested packages from accumulating tools or bypassing denial.

Composition remained visible at the product level.

Skill failure did not corrupt its output contract

A model could stop halfway, produce malformed artifacts, or exceed budget. The skill run returned incomplete with partial evidence and a checkpoint.

Downstream phases consumed only validated artifacts. A person could inspect or resume. The package did not mark success because its final prose sounded complete.

World-state and contract assertions remained independent of the skill's own report.

Expertise participated in evaluation rather than grading itself.

If a package or reference was found unsafe, N17Q disabled selection, removed requested tool paths, and paused active phases at a safe boundary.

Existing effects followed recovery contracts. Accepted artifacts and traces remained readable with the revoked package identity. Resume required another reviewed method and new context transition.

Revocation removed capability without rewriting what had already happened.

Skill lifecycle included an exit.

I removed a skill package and reopened runs that had used it. Their semantic events, artifacts, tools, effects, and evaluations remained readable.

Manual continuation could follow the checkpoint without the packaged instruction. A new agent might need another eligible skill for specialized work.

If deleting a skill erased the meaning of an accepted artifact, its state had lived in the wrong layer.

Skills helped produce work; they did not own it.

References could contain prompt injection. Scripts could be malicious. Templates could leak data. N17Q treated each asset by type and provenance.

Text stayed quoted and scoped. Executable helpers required code review, sandbox execution, explicit arguments, and no ambient credentials or network. Unknown asset types failed.

The model could not invoke arbitrary code because a skill documentation page linked to it.

Packaging convenience never erased the asset boundary.

Privacy influenced selection

A skill might require sending source excerpts to a provider, retrieving external documentation, or retaining diagnostic artifacts. Metadata declared those needs.

The compiler compared them with workspace policy before loading. A local alternative or manual workflow could remain eligible. Context manifests recorded data sent.

The best technical skill could be unavailable for one data class.

Expertise did not override ownership of information.

Loading instructions and references consumed context. Specialized model turns, tools, and artifact generation consumed run resources.

N17Q estimated the activation cost, charged actual usage, and kept recovery reserves. A skill could decline a task that did not fit remaining budget.

The agent could request a bounded extension and could not hide cost inside a nested workflow.

Progressive disclosure controlled resource use as well as attention.

The interface showed active expertise quietly

The run view named active skill and purpose, with access to version, references, requested tools, and limitations. It did not turn every package into an avatar or another conversational participant.

Authors could see why a specialized checklist appeared and dismiss or replace it where policy allowed. Historical skill use stayed in the trace.

The primary workspace and artifact remained visually dominant.

Expertise supported the work without competing to own it.

With the repository skill selected, the agent received applicable project instructions, current local diff, scoped file and test tools, and the procedure for preserving unrelated work.

Replay, MCP registry, export, and grading manuals stayed out of context. Their product controls still existed where relevant.

The run changed one file, ran the correct focused test, preserved the overlay, and produced a compact review artifact.

The model did not become smarter. The product stopped asking it to carry every possible expertise at once.

The trace also became easier to review. Every instruction had a reason to be present, every tool belonged to the active phase, and the resulting artifact named the procedure that shaped it. When the run made a poor choice, I could inspect the relevant package instead of searching a universal prompt for accidental interactions.

That maintainability was as valuable as the saved tokens.

Progressive disclosure is an authority pattern

Skills are often described as a way to save tokens or organize instructions. In N17Q, their deeper value was alignment.

The task earned specialized context. The skill requested precise capabilities. Policy decided what became available. Outputs crossed typed boundaries. The package unloaded when its phase ended. The trace kept the evidence.

This let one agent system handle varied work without one universal prompt and one universal tool set.

Load what the current problem needs.

Keep everything else available, versioned, and out of the way until it has a reason to enter.

The result was not less capable; it was capability arriving with context, scope, evidence, and an accountable exit.

That distinction stayed visible in every trace.