L7 — Operate the harness
Goal. By the end of this lesson you can answer three questions about any run without guessing: what actually happened, what it cost, and how to get the raw evidence back out.
Why here. L6 gave you durable facts. This lesson turns those facts into answers, so that L8's multi-agent fan-out is measurable rather than a leap of faith.
Concepts taught
| Concept | What you learn |
|---|---|
| Session query | ctx.sessionQuery plus the opt-in model-facing session tools |
| Workspace authority | Cross-session access requires exact cwd equality |
| Trajectory replay | Reconstructing a run from the log rather than from a live process |
| Token accounting | dsh-token-meter and dsh-session-stats |
| Telemetry modes | OTel export gated behind explicit user feedback, and how to opt out |
| Invariants | runtime-diagnostics/invariants as a self-check |
| Cost auditing | Turning a trajectory into a defensible number |
Reference: observability & auditing and the telemetry pipeline checklist.
Prerequisites
L1–L6 complete, including the durable state L6 derives — this lesson reads it back. A model provider is required for the cost part; steps 1–4 here are testable without one.
Step 1 — Turn on programmatic session history
Session history access is a service (ctx.sessionQuery) plus, optionally, model-facing tools. Both pieces are opt-in, so enable them deliberately rather than assuming they are present.
The kit ships the overlay at <kit>/solutions/l7.patch.yml. Note its shape: one override and three inserts.
# already mounted by the web bundle, configured never to open — override in place
- id: session-query-sqlite
config:
path: ':memory:'
openAt: startup
# not mounted by any shipped bundle — insert
- insert:
- id: tool-session-query
name: '@deepseek-ai/dsh-tool-session-query'
- id: invariants
name: '@deepseek-ai/dsh-invariants'
- id: session-invariant
name: '@deepseek-ai/dsh-session/invariant'2
3
4
5
6
7
8
9
10
11
12
13
14
The web bundle mounts session-query-sqlite with openAt: never, which provides the capability without opening a store. Turning it on is therefore an override — no insert, no name — the same in-place mechanism you used to override a config in L2 and to disable a provider in L3. The tool package, by contrast, is not in any shipped bundle, so it is an insert.
The tool package must be installed, and pinned
A row naming a package only works if that package is resolvable in the profile. Two things bite here, both hit while building this lesson:
dsh plugin --profile kitdemo add @deepseek-ai/dsh-tool-session-query@<dsh version>- A row naming an uninstalled package fails at load:
tool-session-query (@deepseek-ai/dsh-tool-session-query): failed to import, because a row name resolves through the profile's installation. - Do not omit the version. npm's
latesttag for these optional packages is stale (0.0.1-rc.1), and dsh rejects it outright:Plugin @deepseek-ai/dsh-tool-session-query@0.0.1-rc.1 is incompatible with dsh 0.2.0-rc.2. Install the version matching your dsh runtime —0.2.0-rc.2here — and it installs cleanly and activates.
That second point is a real trap for a reader: the obvious command is the one without a version, and it appears to work until the boot.
Now the model has five read-only tools: session_search, session_event_search, session_trace, session_event_trace, and session_event_read.
Two contract details worth understanding before you rely on them:
- Authorization is workspace-scoped. Cross-session access requires the target session's
cwdto equal the caller's exactly; a caller without acwdcan inspect only itself. Unauthorized boundaries appear as markers without hidden ids, and missing versus cross-workspace guesses behave identically — no information leak either way. - Search is cursor-free and capped. A capped result asks the model to narrow its query rather than exposing offsets or page sizes. The caller's own session is always omitted from search, and for the current session the tools stop before the step that invoked them.
Enabling this package adds fixed guidance plus five tool schemas to every model request, so it is a real prompt-budget decision, not a free toggle.
Two limits of history access, both found by running it
A session the harness cannot interpret poisons search for the whole corpus. Full-text search observes sessions, so one session containing an event type outside the harness vocabulary makes searchSessions fail outright:
session-search persistence observation failed: session "…" contains event type "l6/step"
(seq 4) unknown to this harness and not marked ignorable; refusing to interpret the log2
That is Lesson 6's trap seen from this side — see ADR-0024. readSession still fails for that session; listSessions still lists it.
Retrieval only sees events that carry semantic text — and that is a short list. The query layer builds a searchable document for user and assistant messages, tool calls, tool results, todo writes, and turns that ended with a reason. Every other event contributes an empty string and is dropped ("structural events are omitted"), so a log-only event like sandbox/mode is invisible to filterEvents and to full-text search by design — not because it is invented, and not because anything is broken:
[l7-probe] events the query layer can index (semantic-bearing): 0 of 5
[l7-probe] filterEvents by type 'sandbox/mode': 0 match(es)
[l7-probe] searchSessions: 0 hit(s)
[l7-probe] listSessions: 1269 total; mine found: true
[l7-probe] readSession: 5 event(s); marker present: true2
3
4
5
listSessions and readSession still see the session and its events; only the retrieval layer filters them out, because there is nothing in it to index.
The practical consequence is the one to carry: do not build retrieval on a structural event type, whether you invented it or not. Put the content you want to find in a message, a tool result, or another semantic-bearing event, and it will be found — the turn phase asserts exactly that, with searchSessions(<marker>): 1 hit(s) on a session whose user message carries the marker. A plugin-declared event type has a second, harsher cost, which is the one above: the log stops being readable at all.
Boot with the overlay to confirm the whole composition still activates:
dsh --profile kitdemo --patch <kit>/solutions/l7.patch.yml --port 0 --no-openStep 2 — Read the trajectory instead of the terminal
Terminal scrollback is not evidence. Reconstruct the run from the log:
- Ask the agent to
session_event_readthe events of the session you just ran, or read the log yourself at$DSH_HOME/sessions/<workspace>/session-<uuid>/session.jsonl.zstd($DSH_HOMEis~/.dshby default, and the file is zstd-compressed —zstd -dc <file>). - Find the
sandbox/modeevents L6 folds and confirm their sequence position relative totool/result. - Use
session_event_traceon one event to see its positional replacements and cited source-event relationships. This is how you answer "what did the model actually see at this point" rather than "what does the transcript look like".
The mental model to keep: the session log is the source of truth, deriveMessages() projects model history from it, and the human transcript is a different projection. A trajectory view that cannot be reconstructed from the log is a bug, not a feature.
Step 3 — Account for tokens
Token accounting is always on: token-meter is in the base bundle. Two practical moves:
- Ask for a session statistics view and compare per-turn totals before and after you mount the extra tools from step 1. You have just made a measurable trade: five schemas and fixed guidance on every request, in exchange for retrieval.
session-statsis mounted by the web bundle, so this step needs a web-backed profile —--profile web— not thekitdemoone earlier lessons boot (see ADR-0016). - Force a compaction (
/compact, fromcommand-compact) on a long session and observe the reduction. Compaction is the harness's own answer to context pressure, and watching it once teaches more than reading about it.
Record the numbers. L8 will have you compare a monolithic run against a decomposed one, and that comparison is only meaningful if you can produce these figures on demand.
Step 4 — Know what telemetry leaves the machine
OTel export is mounted by the base bundle but gated: OTel releases a Session-log prefix only after explicit user feedback, regardless of model provider — ordinary activity never triggers capture. Useful controls:
| Control | Effect |
|---|---|
DSH_TELEMETRY_OTLP_URL | Override the production endpoint |
DSH_TELEMETRY_DISABLED | A non-empty value (including '0'/'false') opts the process out |
DSH_TELEMETRY_MODE | The session-telemetry-otel mode, default FEEDBACK_ONLY |
Note the mechanism: launchers patch the row disabled; config alone cannot disable a mounted row. Opacity here is a feature — read the row's own comments in packages/bundle/base/cordis.patch.yml before you change the mode, and decide deliberately whether an exploration workspace should emit anything.
Step 5 — Let the harness audit itself
@deepseek-ai/dsh-invariants checks runtime properties — including the "model-visible means logged" invariant that L6 depended on. It is not mounted by the base, web, or headless bundles (only sdk-minimal carries it), which is why solutions/l7.patch.yml inserts it alongside the session-specific half:
- id: invariants
name: '@deepseek-ai/dsh-invariants'
- id: session-invariant
name: '@deepseek-ai/dsh-session/invariant'2
3
4
The shipped composition pairs them; copying only the first leaves the session checks unarmed. Both are already in the overlay, so a boot with it runs the checks.
A resolution subtlety worth knowing here. Neither package is mounted by any shipped bundle, but unlike the query tool in step 1 they do not need installing: a row naming @deepseek-ai/dsh-invariants activates even in a profile that never installed it. Verified by applying this overlay to a fresh profile, where the invariants rows resolved and only the uninstalled query tool failed — dsh-invariants is not even a dependency of the base bundle. The reason is that rows resolve against the running installation's own package tree, and the lessons require a source checkout, whose workspace provides it.
So the rule is: a row naming a package your dsh installation already contains resolves; a row naming an optional package it does not must be installed first. That is the whole difference between this step and step 1.
An invariant failure here is the cheapest possible way to find a design mistake you would otherwise discover through corrupted replays weeks later.
Then write the audit down. A defensible cost statement has four parts: the task, the turn and step count, the token totals, and the events that prove them. If any part is missing, the number is a guess.
Verification
bash <kit>/solutions/verify-l7.sh <path/to/deepseek-harness>Observable without a session:
- The overlay's shape is right:
session-query-sqliteis an override (the web bundle already mounts it) whiletool-session-queryand the two invariants rows are inserts. openAt: startupis present — the shipped default isnever, so without it the store never opens.- A boot with the overlay produces no activation warnings. The signature of getting this wrong is worth recognizing on sight:
tool-session-query ...: failed to importmeans a row names a package the profile does not have installed. - The profile manifest shows
@deepseek-ai/dsh-tool-session-querypinned to your dsh version — not to npm's stalelatesttag, which dsh rejects as incompatible.
Executed keyless, against the repository's scriptable mock provider (ADR-0027) — a real turn, with the model's output scripted:
A real turn records an assistant answer that full-text search then finds from the trajectory, and the accounting projection is exposed with its documented shape (
totals+last) and real numbers. The mock reportsinput_tokens: 3and an output count equal to its scripted reply's character count, so the check asserts exact non-zero values rather than shape alone./compactsettles as a command ({"kind":"success","text":"No compactable history yet."}on a fresh session).sessionStatsis absent on a base-backed profile, which is why step 3 above tells you to use a web-backed one for statistics.The invariant sweep reports no failure on the kit's composition.
Also executed keyless, but needing a tool call rather than just a turn:
session_event_readreturns one event as JSON, withBefore:/After:neighbour summaries.- A deliberate cross-workspace query is refused, and a missing target is indistinguishable from an unauthorized one.
Still requiring a provider that reports real usage:
- You can state the token delta caused by mounting
tool-session-query. /compactproduces a measurable reduction on a long session.
Items 4–10 are executed and recorded in VERIFIED.md.
Items 9 and 10 need no model, only a scripted tool call: the mock asks for the call and the harness dispatches the tool for real. The refusal in item 10 is produced by the tool executor, so only a dispatched call can reach it — and the check compares the refusals for an existing foreign target and a nonexistent one, because "both were refused" would pass for two different messages.
Item 12 is executed too: after a turn there is history to compact, and /compact reports Compacted 4 history items (~4742 tokens) — a measurable reduction, not just a settled command.
Item 11 is executed too, with a real provider, because it needs a provider whose input tokens reflect the request — the mock reports a constant input_tokens: 3 regardless of what is sent, so no delta can be observed however the request changes. The check holds the composition still and toggles one row: mounting tool-session-query costs 1,664 input tokens (14,544 against 12,880) for the same prompt and model.
Two things about that measurement are worth copying. It requires one step per run, and the first attempt proved why: with a conversational prompt the model took seven steps with the tools mounted and one without, so the comparison measured trajectory rather than schema. And it is opt-in (DSH_REAL_PROVIDER_PATCH), because the repository stays keyless by default. Every item in this lesson is executed, all but this one keyless.
Exit check — you should now be able to explain
- Why the human transcript and the model history are two different projections.
- What the workspace-authority rule prevents, and how it fails closed.
- Why "trajectory replay" is a stronger claim than "we log things".
- Which of the four audit parts you would be tempted to fake, and what makes that detectable.
Next
L8 — Orchestrate multiple agents, where these measurement skills pay off immediately.