I removed the network and found the product
N17Q stopped treating offline execution as a sandbox default and made network absence part of task design, evidence freshness, capability selection, dependency strategy, and the final account.
The agent finished the report offline and cited yesterday's exchange rate as current.
The sandbox had done exactly what its policy promised: no packets left the environment. The model found a cached source file, produced a coherent analysis, and described the result without mentioning that the task's central fact could not be refreshed.
I had treated network denial as a security configuration. N17Q needed to treat it as a change in what the product could honestly accomplish.
Network disabled was not merely how the task ran. It was part of the task's meaning.
Offline changed the evidence horizon
A workspace could contain sources captured at different times, generated datasets, dependency caches, and prior observations. Without network access, those materials defined the outer edge of what the agent could know.
N17Q indexed each resource with origin, captured time, validity window where known, and intended use. Context compilation exposed this provenance. The final claim checker compared time-sensitive statements with the freshest eligible evidence.
The agent could still analyze, transform, draft, and qualify. It could not silently convert available into current.
Network state became an epistemic boundary as well as an execution boundary.
“Research this topic” hid whether historical sources were enough. N17Q added freshness and source requirements to task state: fixed corpus, current as of a date, live verification required, or best available with explicit limitations.
Offline environments were eligible for fixed-corpus work and some artifact tasks. A current-price or latest-policy request paused unless fresh approved material had been staged. The user could narrow the question or provide sources.
The product did not ask the model to infer how much staleness was acceptable from wording alone.
Network eligibility began with what the answer needed to claim.
Denial happened before the shell command
The early sandbox let commands attempt connections and relied on the operating system to fail. Tools waited through timeouts, emitted noisy errors, and tempted the model to try alternate clients or endpoints.
N17Q removed network-dependent product capabilities when the environment was offline. Shell policy denied commands whose declared purpose required external access, while the lower sandbox still blocked unexpected traffic.
The model saw a concise state: offline execution, available staged sources, and safe alternatives. A hidden curl path could not bypass the product decision.
Good denial reduced wasted planning as well as risk.
No network needed an enforceable definition
Disabling a browser tool did not disable DNS, package managers, subprocesses, local proxies, metadata services, or Unix sockets connected to networked helpers.
N17Q's environment contract specified denied outbound interfaces, allowed local fixture services, loopback behavior, inherited descriptors, and host mounts. The sandbox or platform supplied enforcement evidence. Tests attempted common and obscure routes.
Some environments could promise only coarse isolation. Sensitive offline tasks required stronger conformance or did not run there.
The product label described an observed capability boundary, not the absence of one obvious button.
Replay scenarios needed network-shaped behavior without external traffic. N17Q exposed named local fixture endpoints through a separate namespace and policy.
The model invoked product tools, not arbitrary addresses. The trace labeled responses simulated. DNS and connection rules prevented fixture names from resolving outside the environment. A missing fixture failed closed.
This allowed realistic retries, delays, and streaming while preserving the offline guarantee.
The interface distinguished network disabled from no socket calls at all because the implementation and evidence differed.
Package installation became preparation
Agents often discovered a missing dependency halfway through a task. With network disabled, install either failed or reached an undeclared cache whose contents varied.
N17Q resolved skill and environment requirements before execution. Eligible packages came from a sealed image or content-addressed store with lockfile and provenance. The model could use them but could not expand the dependency graph silently.
If a new library was genuinely needed, the run paused and prepared a separate dependency acquisition step under network policy. The acquired environment received a new manifest and affected verification became stale.
Offline work became reproducible because dependency needs were treated as inputs.
Cached data was an observation with an age
A browser cache, downloaded API response, or previous search result could be useful. Calling it “cache” did not explain whether it was trustworthy or appropriate.
N17Q stored capture method, source URL or handle, retrieval time, response metadata, content digest, and policy. The context compiler could include a cached resource only if the task allowed that purpose and age.
A stale source might support historical comparison while being disallowed for a present-tense claim. The final report named the cutoff.
Offline execution did not make cached data wrong. It made its limitations impossible to refresh away.
Some workflows needed one official source or one package registry, not the open internet. A boolean network flag forced a choice between incapability and excessive reach.
N17Q modeled approved connections and destinations. Environment controls enforced allowlists or product proxies where supported. Secrets were scoped to destination and kept outside model-visible state. Reads and effects still passed semantic policy.
The task could therefore move from fully offline to one bounded acquisition phase, then return to an isolated artifact phase.
Network access became a set of explicit capabilities with purpose and evidence.
Fetching public documentation and sending a private artifact both used the network. Their effect and data semantics were entirely different.
N17Q separated acquisition, observation, and external mutation. Research connections allowed bounded reads and recorded sources. Delivery required materialized artifact, destination, approval, effect identity, and recovery contract.
Enabling research never made upload or messaging available. A domain used for reading could not receive workspace data through an arbitrary shell request.
One network switch could not express the difference, so the product vocabulary had to.
Data egress mattered more than request direction
A GET request could include sensitive query parameters, headers, path fragments, or DNS names. A package install could transmit environment metadata. Calling the operation read-only said nothing about what left the boundary.
N17Q's connection contracts declared outbound data classes and credential behavior. The network proxy enforced destination; product policy inspected normalized request purpose and payload. Logs stored safe projections.
Offline mode provided the strongest egress restriction. Bounded online modes explained exactly which information could cross.
Security review moved from HTTP verbs to actual data flow.
Tool output sometimes suggested disabling SSL checks, using a mirror, switching ports, or trying another hostname. A model following those instructions could broaden reach through persistence.
N17Q kept network configuration outside model authority. The agent could propose that fresh data was needed and name the evidence gap. A person or product workflow could approve a connection change after reviewing destination and data.
The current run did not gain connectivity from a shell flag, environment variable, or persuasive error message.
Network policy changed through explicit state, not troubleshooting improvisation.
Isolation was not only a restriction. It reduced nondeterminism, protected fixed-corpus analysis from source drift, eliminated accidental disclosure, and made fixture replay dependable.
N17Q used offline execution for document transformation, repository analysis, deterministic evaluation, and synthetic scenarios. Inputs could be sealed, outputs compared, and network absence asserted.
The interface described this as a property of the run rather than a degraded mode. When freshness was irrelevant, offline was often the better environment.
Product design improved when capability was selected from task needs instead of always maximized.
Offline did not mean trusted input
A malicious document remained malicious after download. A staged repository could include scripts that attempted to read protected paths. Network denial reduced exfiltration routes and did not make instructions or code safe.
N17Q preserved provenance, separated instructions from content, sandboxed execution, constrained mounts, and scanned candidate artifacts. Tool results remained observations.
The product avoided presenting offline as a universal safety badge. It named the threats the boundary addressed and the ones other controls handled.
Honest security claims were narrower and easier to test.
A small crossed-globe icon told users the run was offline and not what that changed. N17Q displayed source cutoff, unavailable capabilities, staged evidence, and any task requirement that could not be met.
For the report scenario, the header said: Offline analysis — sources current through 27 April 2026; live rate verification unavailable. The final answer inherited that limitation.
If the user enabled one research connection, the view showed the exact destination and acquisition purpose. Delivery remained absent.
Network state became intelligible product context rather than a hidden sandbox detail.
Final claims were checked against connection evidence
The phrase “I checked the current rate” implied a live observation. N17Q extracted that claim and looked for an eligible source event within the task's freshness window.
An old file did not satisfy it. A fixture response could satisfy a simulated benchmark and was labeled accordingly. A live source needed connection receipt, retrieval time, and content provenance.
The system could not judge every subtle temporal phrase deterministically, so a model grader reviewed remaining claims with the evidence package.
The product made unsupported freshness difficult to hide behind fluent tense.
Network failures did not become offline mode silently
If an approved source timed out, falling back to cache could be useful. Doing so without changing the task account would misrepresent the evidence.
N17Q recorded connection failure, evaluated whether fallback met the freshness contract, and either continued with an explicit degraded state or paused. The model received the change as structured context.
Retry budgets and provider status informed the decision. Alternate destinations required their own eligibility.
Offline was an explicit outcome with consequences, not the accidental result of a broken network.
Long-running tasks revalidated network policy
A run could pause overnight while allowlists, connection permissions, or data classification changed. Yesterday's connectivity could not be assumed.
At each acquisition or effect boundary, N17Q evaluated current policy and environment evidence. A revoked connection disappeared from the catalogue. Existing observations remained in the trace with their original provenance; they did not become retroactively unread.
Pending work could continue locally or stop with a specific freshness gap.
Durable state allowed network authority to change without making the run forget what had already happened.
An offline process attempting a denied connection was important evidence. Calling it a network effect would overstate impact; omitting it would hide behavior.
N17Q recorded command proposal, product denial where available, sandbox denial, destination projection, and zero successful connection receipt. Repeated attempts affected progress and agent evaluation.
Reports could therefore say both that containment held and that the proposed behavior required attention.
World integrity and model quality remained separate findings.
Claiming no network activity required more than an empty application log. N17Q fixture environments ran with denied routes, empty live connection registry, no credentials, and conformance probes. The supervisor accounted for subprocesses and background tasks.
The invariant asserted no successful external connection through the declared environment boundary. It did not claim visibility into hardware or infrastructure outside that contract.
When an environment could not provide sufficient evidence, the result became indeterminate and high-sensitivity tasks stayed ineligible.
Negative claims earned trust by naming where enforcement and observation occurred.
Artifact import was the offline ingress path
If a task needed fresh material, one option was to acquire it outside the execution environment and stage it deliberately. That boundary needed its own controls.
N17Q imported files through a product flow that recorded source, capture time, content digest, declared purpose, scanner results, and person or capability responsible. Archives were expanded under limits. Active formats received safe previews. The resulting snapshot became a new task input revision.
The agent could not watch a download folder and absorb whatever appeared. New material entered at a checkpoint, invalidated affected conclusions, and updated the evidence horizon visibly.
Offline work stayed sealed while still allowing accountable refresh.
Even when outbound connections failed, DNS queries could reveal internal project names or sensitive host fragments. A command attempting a unique domain could communicate through the lookup itself.
N17Q offline environments denied external name resolution and used fixed local mappings only for fixture services. Connection-capable environments resolved approved destinations through the controlled network layer rather than exposing a general resolver to arbitrary subprocesses.
Conformance tests inspected resolution as well as completed sockets. Logs retained safe destination classifications without copying secret-bearing hostnames into normal views.
Egress policy began before an HTTP request existed.
Application metrics, crash reporters, package analytics, and shell history synchronization could create network traffic unrelated to the agent's apparent tools.
N17Q disabled or redirected telemetry inside offline environments and documented the platform services still required for orchestration. Where a hosted control plane necessarily exchanged command and result data, the product did not describe the environment as physically disconnected. It described external destination access as denied and named the trusted platform boundary.
This distinction prevented a strong-looking label from hiding necessary infrastructure communication.
Network language became precise enough to survive technical review.
Offline outputs needed a release boundary
An isolated run could generate an artifact containing sensitive or malicious content. Moving it back to the user or another system was itself a data transfer, even if the agent had no general network.
N17Q sealed outputs, scanned them under policy, rendered safe previews, and required explicit selection for export. The release receipt named artifact digest and destination boundary. Automated external delivery remained a separate effect.
Discarded or quarantined outputs never left through convenience download links. The run report distinguished created inside the sandbox from released outside it.
Isolation protected the work phase; controlled export protected its exit.
Two runs with identical source and model could diverge because one resolved a live dependency, fetched a changing page, or reached telemetry that altered timing. Omitting network policy from the environment manifest made comparisons incomplete.
N17Q recorded connection catalogue digest, enforcement mode, allowed fixture endpoints, observed acquisition receipts, and unexpected denied attempts. replay runs replaced live sources with captured fixtures or stopped on missing evidence.
A report could therefore explain whether divergence came from intelligence, source freshness, connection availability, or environment drift.
Reproducibility improved when network became an input, not atmospheric background.
The exchange-rate report became honest
N17Q reran the task in an offline snapshot containing data captured the previous afternoon. The context compiler surfaced the timestamp. The agent produced a historical comparison and labeled the rate as last observed, not current.
The source panel showed exactly which figures came from the sealed dataset and which conclusions were derived locally. It did not display a freshness checkmark whose meaning depended on network availability. If the artifact was reopened later, the same cutoff remained attached instead of being silently reinterpreted relative to the new date.
Because the user required a current figure, the run stopped before final delivery and offered one narrow acquisition step from an approved source. In a second branch, the user accepted an as-of report; the artifact and final answer carried the cutoff.
That design made offline output durable. A reader did not need access to the original execution environment to understand the evidence horizon, and a future refresh could compare a new acquisition against a known prior snapshot rather than replacing it without lineage.
Both branches were useful. Neither pretended isolation had no effect on the answer.
Disabling network can protect data, stabilize execution, and reduce the agent's attack surface. It can also remove the evidence, dependencies, or destinations a task requires.
Make that trade in product state. Let it shape planning, claims, interface, and evaluation. The sandbox should enforce the choice—but the product must explain what the choice means.
Only then can offline work be both safe and intellectually honest about its limits.
By design, every time.
Without exceptions.