Kiwi runtime audit

Audit of src/kiwi_runtime/main.py on branch AUT-1688-disconnect-issue-p1, 2026-10-02. The file is 5,960 lines and 151 functions. I read all of it and ran an experiment for each suspected problem. Nothing in this report has been changed in the code yet.

How to read it:

The work order agreed is: refactor first (section 1), then the bugs (section 2), then the event loop, parallelism, memory and computation (sections 3 to 5).

Refactor status

Section 1 is done on this branch, uncommitted. No behaviour was changed; the bugs in sections 2 to 5 are all still present and now each sits in a small file.

Before After
Files 1 (5,960 lines) 27, the largest 659 lines
Longest function 2,165 lines 331 lines (search_in_files, see below)
_connect_once 456 lines 56 lines
main() 266 lines 38 lines
Copies of the binary-file response 10 2 (one for chunk mode, whose message differs)
_strip_line_endings definitions 4 1

How it was checked:

Where the old line ranges went:

Old lines Module
97-331, 3480-3693 auth.py
371-419, 487-659 fs/text.py
662-703, 3226-3476 restricted.py (_looks_interactive_oneoff, 3319-3390, is in commands.py)
707-840 checkpoints.py
842-3033 fs/: read.py, write.py, paths.py, search.py, upload.py, batch.py, codelens.py, handler.py
3062-3221 ui.py
3746-3974 commands.py
3035-3052, 3977-4572 sessions.py
4575-4810 tui_bridge.py
4813-4961 uploads.py
4964-5099 state.py
5102-5687 link.py (connect, authenticate, receive loop) and handlers.py (one per message)
3055-3059, 3696-3743, 5690-5955 cli.py

Differences from the plan:

Fix status

Step Items State
1 Refactor section 1 done
2 PTY sessions RT-01, RT-02, RT-03, RT-04, RT-21 done
3 Remaining bugs RT-05 to RT-09 done
4 Event loop RT-10 to RT-14 done
5 Parallel commands, bounded output RT-16, RT-20 done
6 Memory and speed of file operations RT-19, RT-22 to RT-27, RT-29 done
7 Process pool RT-17, RT-28 done

Step 2 as built (sessions.py, state.py, handlers.py, link.py):

Not done from the proposal: the backend does not send pty_close (it has no signal that a run is over, so the runtime's own limits do the work), and file requests still share the default thread pool, which no session occupies any more. PipeProcess (Windows) is unchanged.

Checked by ten new tests in tests/test_runtime_connection.py, each against a real shell: a 20,000-character command, a heredoc, shell state across commands, the script file's removal, 50,000 non-ASCII characters intact, a 2,000,000-character line, a file request answered with 16 sessions open, eviction, release on exit, and the idle close. Suite: 220 passed.

Step 3 as built:

Checked by tests/test_runtime_fixes.py (59 cases) and two more in tests/test_runtime_connection.py. Suite: 281 passed. The 3,345 recorded read and search results are unchanged.

Steps 4 to 7 as built:

Measured after, longest time the event loop was held:

Work Before After
Printing to an output nobody reads as long as the stall (1,996 ms measured) not held (test bound 200 ms)
Token lock held by another program for 1 s 1,070 ms not held (test bound 200 ms)
read_file, 1,000,000 characters of a 60 MB file 189 ms 9 ms
One-off command printing 200 MB 400 MB peak memory 1.4 ms held, under 10 MB
Four sleep 0.5 commands sent together 2.05 s in total 0.52 s in total
Search of 1,332 files for a rare literal, in the worker 126 ms 63 ms
Eight regex searches at once shared the interpreter with the loop 6 ms held

The first search after start takes about 0.4 s longer, while the worker processes start. The 9 MB frame still holds the loop for the time it takes to parse (about 100 ms).

Checked by 16 more tests. Suite: 297 passed. The 3,345 recorded read and search results are unchanged, also with symlinks inside and outside the allowed directory in the tree.

Summary

Count Worst case
Bugs, confirmed 9 Sessions are never closed and starve all file requests (RT-01)
Event-loop blockers 6 Loop frozen for as long as another process holds the token lock (RT-10)
Parallelism gaps 3 One-off commands run strictly one at a time (RT-16)
Memory 5 400 MB to return 200,000 characters (RT-19, RT-20)
Computation 6 Search 13 times slower than it needs to be (RT-24)
Structure 1 file One function of 2,165 lines

Evidence comes from two scratch scripts, runtime_audit.py and search_bench.py. They run the real connect loop against a local WebSocket server with real shells, or call the file handler directly. They are not part of the repository.


1. Refactoring plan (first)

Why first

Every fix below lands in one of two giant functions. Fixing inside them makes the diffs hard to review and the next reader's job harder. Splitting first means each later fix is a small change in a small file.

The risk of refactoring first is breaking behaviour while moving code. The plan below keeps that risk low: moves are mechanical, each step is one commit, and the 205 existing tests must pass after every step.

What is wrong with the structure today

Fact Figure
Lines in one file 5,960
Longest function, _handle_fs_request_sync 2,165 lines, 28 operations in one if chain
Second longest, _connect_once 456 lines, 12 message types in one if/elif chain
main() 266 lines: argument parsing, login and the run loop together
Copies of the "File appears to be binary" response 10
Definitions of _strip_line_endings 4, identical, nested in different branches
os.killpg call sites for "kill the command's process group" 7
import statements inside functions 32
The grep-with-context loop written twice (read_file grep mode and search_in_files)

Unrelated concerns share the file: authentication, file operations, checkpoints, restricted-mode validation, command execution, PTY sessions, the connection, terminal output and the command line.

Target layout

src/kiwi_runtime/
  main.py            entry point; re-exports the names tests and the TUI import
  cli.py             argument parsing, login menu, run loop        (~350 lines)
  auth.py            tokens.json, refresh, login, OAuth            (~500)
  ui.py              colours, banner, print_* helpers              (~200)
  restricted.py      validate_fs_path, validate_command            (~300)
  checkpoints.py     snapshot before mutation                      (~150)
  commands.py        one-off and interactive commands              (~400)
  pty.py             PTYProcess, PipeProcess, _CommandTracker      (~700)
  tui_bridge.py      _LocalInteractivePTYBridge                    (~250)
  link.py            RuntimeLink, _RequestRunner, connect loop     (~450)
  handlers.py        one function per inbound message type         (~450)
  fs/
    __init__.py      dispatch table, handle_fs_request             (~120)
    context.py       FsRequest: mode, allowed_dirs, request_id     (~80)
    text.py          line endings, encoding, binary detection      (~150)
    read.py          read_file: full, range, head, tail, chunk, grep (~600)
    write.py         write, append, replace, line edits, diff      (~450)
    paths.py         stat, list, mkdir, move, copy, delete         (~350)
    search.py        search_in_file, search_in_files               (~400)
    upload.py        chunked writes                                (~200)
    batch.py         batch                                         (~170)
    codelens.py      the eight code-analysis operations            (~220)

The two structural changes

A dispatch table instead of an if chain. Each operation becomes one function with one signature, registered by name:

FS_HANDLERS: dict[str, Callable[[FsRequest], dict[str, Any]]] = {
    "read_file": read.read_file,
    "write_file": write.write_file,
    ...
}

def handle(request: FsRequest) -> dict[str, Any]:
    handler = FS_HANDLERS.get(request.operation)
    if handler is None:
        return request.failure(f"Unsupported file operation type: {request.operation}")
    try:
        return handler(request)
    except Exception as e:
        return request.failure(str(e))

FsRequest carries what every branch rebuilds today: the message, mode, allowed_dirs, request_id, and helpers resolve(path), ok(**fields) and failure(message). That removes the repeated "validate path, return failure dict" preamble from all 28 operations.

The same for inbound messages. _connect_once keeps only connect, auth and the receive loop. Each message type becomes a function in handlers.py taking the runtime state and the message, looked up in a table.

Steps, one commit each

  1. Pure moves. Create the modules and move functions unchanged. main.py re-exports every name the tests and kiwi_tui import (_handle_fs_request_sync, apply_unified_diff_to_text, RuntimeLink, _CommandTracker, _LocalInteractivePTYBridge, _tui_managed_control_dir, connect, login, the constants). No logic changes. Tests pass.
  2. Split the file handler. One function per operation, the dispatch table, FsRequest. The body of each branch moves as is.
  3. Remove the duplication the split exposes. One binary-file response, one _strip_line_endings, one grep-with-context function used by both search paths, one "kill the process group" helper.
  4. Split the message loop into handlers.py with its table.
  5. Split main() into parse, authenticate and run.
  6. Move the imports to module level, except the platform-specific ones.

Rules while doing it:

Expected result: no file over about 700 lines, no function over about 150, and each fix in sections 2 to 5 touches one small file.


2. Bugs

Severity: critical = the runtime stops doing its job; high = wrong results in normal use; medium = wrong results in specific cases; low = cosmetic or rare.

PTY sessions

RT-01 Sessions are never closed and starve file requestscritical

What happens:

  1. Each PTY session reads its output with loop.run_in_executor(None, self._blocking_read) (:4075). _blocking_read loops on select until the shell exits (:4136-4156), so every session holds one thread of the default thread pool for its whole life.
  2. File requests run through asyncio.to_thread (:3018), which uses the same pool.
  3. The pool has min(32, cpu_count + 4) threads.
  4. Nothing closes a session. The backend no longer sends pty_close, and a session whose shell exited is never removed from pty_sessions either, so its file descriptor stays open.

Evidence: on this machine the pool has 16 threads. With 16 idle sessions open, a stat_file request got no answer within 5 seconds.

Result: a runtime that has opened enough sessions stops answering every file request, and each new session cannot read its own output. Each leaked session is also a live shell process.

Solution:

RT-02 A session command longer than about 4 KB never finisheshigh

What happens: send_input writes the whole command with one os.write on a non-blocking descriptor and ignores the result (:4158-4161). In canonical mode the kernel's terminal line buffer holds 4,096 bytes, so a longer line is cut and the shell never sees its newline.

Evidence: a 3,012-byte command line finished with exit 0. A 5,012-byte one never finished.

Result: a long command (a long python -c, a long path list) hangs until the tool's timeout, and the model is told to avoid heredocs, which was not the cause.

Solution: do not type the command. Write it to a private temp file and type one short line that sources it:

def wrap(self, command: str) -> str:
    script = self._write_script(command)          # 0600, under ~/.kiwi/tmp
    return f". {shlex.quote(script)}; printf '\n{self._mark}%s__\n' \"$?\"; rm -f {shlex.quote(script)}\n"

Sourcing runs in the session's own shell, so cd and export still persist. It removes the length limit, and multi-line commands and heredocs work, because the shell reads them from a file instead of a terminal. The typed line stays under 200 bytes. Restricted-mode validation still runs on the original command text.

RT-03 Non-ASCII output is corruptedhigh

What happens: _blocking_read reads 4,096 bytes and decodes each piece alone with errors="replace" (:4143-4146). A character split across two reads becomes two replacement characters.

Evidence: of 50,000 é printed in a session, 49,988 arrived intact and 24 replacement characters were inserted.

Solution: one incremental decoder per session, codecs.getincrementaldecoder("utf-8")("replace"), which keeps the incomplete bytes for the next read. PipeProcess has the same pattern (:4390).

RT-04 Starting: is printed twice for every session (low)

pty_start logs the same line at :5486 and :5489. Remove one.

File operations

RT-05 replace_in_file with an empty old_text corrupts the file (high)

What happens: content.count("") is the length plus one, and with replace_all Python inserts the replacement between every character (:1692, :1721).

Evidence: on a file containing ab\n, the call reported success: true and left XaXbX\nX.

Solution: reject an empty old_text before reading the file.

RT-06 read_file chunk mode rejects non-ASCII files as binary (medium)

What happens: the chunk is read at an arbitrary byte offset and decoded strictly (:1396-1404). A chunk that starts or ends inside a multi-byte character fails to decode and is reported as binary.

Evidence: a file of 100 é, read with start_byte=1, returned "File appears to be binary in the requested byte range".

Solution: trim the edges instead of failing. Skip leading continuation bytes (0x80-0xBF), drop an incomplete trailing sequence, and report the byte range actually returned so the next chunk starts at the right place.

RT-07 The "interactive" guess refuses ordinary commandsmedium

What happens: any command whose first word is in hard_interactive (:3368-3373) is refused, whatever its arguments.

Evidence: all of these were refused as interactive: psql -c "select 1", ssh host uptime, sudo -n true, mysql -e 'select 1', redis-cli get k, less -F f.

Solution: one-off commands already run with stdin set to /dev/null (:3768), so a tool that wants a terminal fails at once instead of hanging. Keep the refusal only for full-screen programs (vim, vi, nano, top, htop, less, more, watch) and for tools started with no arguments at all (python, node, psql, mysql).

Authentication

RT-08 A timezone-aware expires_at makes connecting impossible (medium)

What happens: _parse_expires_at turns a Z suffix into an aware datetime (:117-118). _token_is_expired then compares it with naive datetime.now() and with the naive JWT expiry (:145-151).

Evidence: TypeError: can't compare offset-naive and offset-aware datetimes.

Result: the exception leaves _get_valid_access_token, which runs at the start of every connection attempt, so the runtime would retry forever without connecting. I did not check whether any writer of tokens.json produces that format today; the parser accepts it on purpose.

Solution: use aware UTC everywhere in this code: datetime.now(timezone.utc), datetime.fromtimestamp(exp, timezone.utc), and treat a naive stored value as local time.

RT-09 Smaller itemslow


3. What blocks the event loop

The event loop is one thread. While anything blocks it, nothing else runs: no incoming message, no keepalive, no keystroke, no output.

RT-10 The token file lock is taken with a blocking callhigh

What happens: _TokensLock.__enter__ calls fcntl.flock(..., LOCK_EX) (:198), which waits as long as another process holds the lock. _get_valid_access_token takes it on the loop (:276) and then awaits an HTTP refresh of up to 30 seconds while still holding it (:285). _load_tokens_file also sleeps with time.sleep on retry (:165).

Evidence: while another process held tokens.lock, the event loop was frozen for 1.07s. It freezes for as long as the other process holds the lock.

When it runs: at the start of every connection attempt and for every read_files upload. kiwi-code and the runtime share this lock, so it is contended exactly when both refresh an expired token.

Solution: run the whole critical section in a worker thread with a synchronous HTTP call. The lock exists to stop two processes refreshing the same refresh token, so it must stay held across the refresh; it just must not be held on the loop.

async def _get_valid_access_token(http_base_url: str, fallback_token: str) -> str:
    return await asyncio.to_thread(_valid_access_token_sync, http_base_url, fallback_token)

RT-11 read_files reads files from disk on the loop (medium, not measured)

handle_read_files opens each file and hands the handle to httpx (:4893-4901). httpx reads it in pieces synchronously while sending. A large file on a slow disk stalls the loop repeatedly for the length of the upload. Solution: stream the file through an async generator that reads each piece in a thread.

RT-12 Printing happens on the loopmediumnot measured

Every line of session output is printed with print from the read loop (:4083-4085). print blocks when stdout cannot accept more: a paused terminal, or a pipe whose reader has fallen behind (the TUI runs the runtime as a subprocess). I did not test it with the TUI. Solution: a bounded queue drained by one writer thread; drop log lines when it is full.

RT-13 The TUI interactive bridge does file I/O per chunk on the looplownot measured

append_output opens, writes and closes output.log for every output chunk (:4628-4633), _write_status reads and rewrites a JSON file, and pump_inputs polls a file every 0.1s. Solution: keep the file open, and move the writes to a thread.

RT-14 Restricted-mode validation does disk and process lookups on the looplow

For each session command, validate_command asks psutil for the shell's working directory and resolves every path in the command (:3268-3317). It is small, but it runs on the loop for every keystroke frame. Solution: asyncio.to_thread.

RT-15 Large messages are encoded and decoded on the looplownot measured

json.loads of an inbound frame of up to 10 MiB and json.dumps of a 1,000,000-character read_file reply each take tens of milliseconds. Acceptable; noted for completeness.


4. Parallelism

RT-16 One-off commands run strictly one at a timehigh

What happens: a single _command_worker takes commands from the queue and awaits each to completion (:5118-5143).

Evidence: four independent sleep 0.5 commands sent together took 2.05 seconds.

Result: one slow command delays every other one-off command, and their timeouts are already counting on the backend. A run with subagents shares one runtime, so this is common.

Solution: one-off commands share no state (each starts a fresh shell), so run up to four at once. Four workers on the same queue is the smallest change:

workers = [asyncio.create_task(_command_worker(state, mode, allowed_dirs)) for _ in range(4)]

running_command, used only by the debug heartbeat, becomes a count.

RT-17 CPU-bound requests run in threadsmediumnot measured

search_in_files, search_in_file and the eight code-analysis operations are pure Python loops. In threads they compete with the event loop for the interpreter lock, so up to eight of them at once slow keepalives, keystrokes and output. Solution: run those operations in a ProcessPoolExecutor with two workers. They take only plain data and return plain data, so they move without change. Measure loop latency before and after.

RT-18 File requests that change files run one at a timeby design

This is the ordering lane added in this branch, so chunked uploads and back-to-back edits stay correct. It could be relaxed to one lane per file path later. Not worth doing until there is a measured need.


5. Memory and computation

Memory

RT-19 read_file full mode loads the whole file (high)

What happens: _read_text_file_with_binary_check reads and decodes the entire file (:376-385, called at :943), then keeps 200,000 characters (:964).

Evidence: reading a 200 MB file peaked at 400 MB of Python memory to return 200,000 characters.

Solution: read only what is returned. Read max_chars * 4 bytes, decode, cut to max_chars. Count the file's lines by scanning it in 1 MiB blocks without keeping them, or omit total_lines above a size threshold as tail mode already does (:1295).

RT-20 A one-off command's output is held entirely in memoryhigh

What happens: proc.communicate() collects all of stdout and stderr (:3775); the result is then cut to 51,200 characters (:3778).

Evidence: a command printing 200 MB peaked at 400 MB; 51,200 characters were returned.

Two more defects in the same lines:

Solution: read both streams incrementally into a bounded buffer that keeps the first 10 KB and the last 40 KB, with a line saying how much was dropped. Memory is then constant, the tail is kept, and a timeout returns what was printed.

RT-21 The session log buffer has no limitmedium

What happens: output is appended to _log_buffer and only released at a newline (:4082-4085). Each chunk also rescans the whole buffer for a newline.

Evidence: after a 3,000,000-character line with no newline, the buffer held all 3,000,000 characters. A progress bar that redraws with \r, or base64 output, does this.

Solution: flush the buffer when it passes 8 KB, and search only the new chunk for the newline.

RT-22 Checkpoints read each file wholemediumnot measured

_checkpoint_before_mutation snapshots with path.read_bytes() (:819). Deleting a directory snapshots every file in it this way. Solution: copy in blocks (shutil.copyfileobj) straight to the snapshot file.

RT-23 tail mode builds its buffer quadratically (low)

buf = chunk + buf and buf.count(b"\n") on the whole buffer for every 8 KB block (:1240-1249). Solution: keep the blocks in a list and count newlines in the new block only.

Computation

RT-24 search_in_files scans every line of every file in Python (high)

What happens: every file is read line by line, each line decoded and tested (:2487-2565). Most files contain no match at all.

Evidence, over 1,332 Python files:

Query Today Whole-file check first
a rare literal (2 matches in 1 file) 246 ms 19 ms

Solution: read the file once (it is already capped at 2 MB), test needle in data on the bytes, and run the per-line loop only on files that contain it. For a regex, compile it for bytes and use search on the whole file first. That is about 13 times faster for the common case of a specific identifier. A search that stops early on its match limit is unaffected.

Two smaller costs in the same loop:

RT-25 list_dir makes three to four system calls per entry (medium, not measured)

_list_entry calls os.stat, os.path.isdir and os.path.isfile (:436-449). Solution: os.scandir, whose entries carry type and size from one call.

RT-26 Checkpoint metadata is rebuilt for every filemediumnot measured

For each file, _checkpoint_before_mutation calls ensure_run_layout, ensure_meta and get_active_entry_id again (:781-790), all of which touch the disk. A directory operation repeats that per file. Solution: resolve them once per request and pass them in.

RT-27 A new HTTP client for every requestlow

_refresh_tokens, login and handle_read_files each create an httpx.AsyncClient (:228, :3490, :4890), which means a new TLS handshake each time. Solution: one client for the life of the runtime, closed at shutdown.

RT-28 A new CodeLens for every request (not measured)

All eight code-analysis operations construct CodeLens() (:2807 and seven more). If the object caches parsed files, that cache is thrown away each time. I did not read pylens, so this may be free. Check before changing.

RT-29 search_in_file loads the whole file and all its lines (low)

read_text().splitlines() (:2065) for a single-file search. It also still splits on form feeds, unlike the edit operations fixed in this branch. Solution: share the streaming grep function that step 3 of the refactor creates.


6. Order of work

Step What Items
1 Refactor, six commits, no behaviour change section 1
2 PTY sessions: event-loop reader, lifecycle, sourced commands, incremental decoder RT-01, RT-02, RT-03, RT-04, RT-21
3 Remaining bugs RT-05, RT-06, RT-07, RT-08, RT-09
4 Event loop: token lock, uploads, printing RT-10 to RT-14
5 Parallel one-off commands and bounded output RT-16, RT-20
6 Memory and speed of file operations RT-19, RT-22 to RT-27, RT-29
7 Process pool for CPU-bound requests, after measuring RT-17, RT-28

Each step ends with the test suite green and, for steps 2 and 5, the end-to-end check with the real runtime process.

7. Limits of this audit