Skip to content

Feature: let a host observe shell commands as they run - #95

Merged
DevMando merged 1 commit into
mainfrom
feature/agent-command-output-sink
Sep 10, 2026
Merged

DevMando merged 1 commit into
mainfrom
feature/agent-command-output-sink

Conversation

@DevMando

Copy link
Copy Markdown
Owner

Summary

Adds an optional hook that lets a host application watch the assistant's shell commands as they run — the command, its output line by line, and how it ended. Nothing uses it yet in the CLI; it exists so MandoCode Desktop can show the user a live view of commands instead of only reporting them afterward.

Why this matters

When the assistant runs a shell command today — a build, a test run, git status — the user sees a single truncated status line while it works, then the finished output. A dotnet build that takes ninety seconds is ninety seconds of near-silence. There is no way for the host to show the work in progress, because the information needed to do that is discarded as it is produced.

Everything a live display needs already exists inside the command runner. It is simply not offered to anyone. This makes it available.

What is new

  • An optional command output sink. A host can attach one and be told when a command starts (with its working directory), each line of output as it arrives with stdout and stderr distinguished, and how the command ended — either the exit code or the reason it was killed.
  • It is observational, and enforced as such. The sink cannot block, alter, delay, or cancel a command. Every callback into it is wrapped, so a fault in a host's display is swallowed rather than propagating into the command the assistant is running. A test covers this directly by attaching a sink that throws on every call and confirming the command still runs and reports normally.
  • It sees more than the model does. Output sent to the model is capped at 5000 characters to protect its context window. The sink is deliberately fed past that cap, because a scrolling display has its own limits and inheriting the model's would hide the tail of exactly the long output someone is watching.
  • A bare cd is not reported. That case is intercepted and never starts a process, so announcing it would leave a command header on a display with nothing under it.

Scope and risk

Low. One new optional parameter, defaulted to nothing, on two constructors. Callers that pass no sink — which is every caller today, including the CLI — take an identical path to before. No change to what the model receives, to timeouts, to output caps, or to how commands are launched or killed.

The one thing worth a reviewer's attention: output callbacks are raised on the command's output-reader threads, so an implementation must be thread-safe and fast, since each callback runs inline with reading the command's output. This is documented on the interface, and the accompanying Desktop change is expected to hand off to its UI thread immediately.

Verification

Full engine suite passes on both target frameworks — 733 of 733 on .NET 10 and .NET 8. Five new tests cover the lifecycle end to end: start/output/exit reporting, a non-zero exit reported as a failure rather than a kill, a bare cd producing no report, output surviving past the model's truncation cap, and a throwing sink leaving the command unaffected.

Not covered: no test exercises a killed command's reporting path, since that requires waiting out the 30-second idle timeout. The code path is one line alongside the existing kill handling.

Dependency

Nothing depends on this landing first, but it is a prerequisite for the MandoCode Desktop change that adds a live agent output view. That work needs this merged and the Desktop submodule pin moved forward.

execute_command already produces everything a live display needs — the
command, its working directory, each line of output as it arrives, and an
exit code — and then throws all of it away except a one-line spinner
activity string truncated to 80 characters.

Add an optional ICommandOutputSink a host can attach to receive that
lifecycle. Purely observational: it cannot gate or alter a command, and
every callback is wrapped so a faulting display can never propagate into the
command the agent is running.

Deliberately teed OUTSIDE the 5000-character output cap. That cap protects
the model's context window; a host display has its own scrollback, and
inheriting the cap would hide the tail of exactly the long build output
someone opened a panel to watch.

Nothing changes for a caller that attaches no sink — the CLI passes none and
behaves as before. The sink is handed to the plugin on every agent rebuild,
so a host attaches once and keeps it across model switches and settings
changes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant