fix(llmfs): expand model aliases and update default model ID
Add a modelAliases map in session_settings.go so that short names
written to /n/llm/N/model are expanded to full Anthropic model IDs:
haiku → claude-haiku-4-5-20251001
sonnet → claude-sonnet-4-6
opus → claude-opus-4-6
This fixes spawn subagents which write short names like "haiku" to the
model file — the Anthropic API rejects bare aliases as invalid model IDs.
Also update the default model from the stale claude-sonnet-4-20250514
to claude-sonnet-4-6 to match current API availability.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
feat(llmfs): async Write + per-session streaming for live token delivery
Enable lucibridge to stream LLM tokens into the Lucifer conversation
zone as they are generated, without waiting for the full response.
Changes:
- session.go: add BeginGeneration/EndGeneration/SendChunk/GetStreamCh/
WaitDone methods on Session; streamCh (cap 256) carries raw text chunks
during generation; doneCh signals completion; EndGeneration closes
channels but does NOT nil streamCh (late readers still see closed chan)
- session_ask.go: Write() is now async — calls BeginGeneration(), spawns
goroutine, returns immediately; Read() calls WaitDone() before accessing
LastResponse so pread blocks until generation completes
- session_stream.go (new): /n/llm/N/stream file; Read() blocks on <-ch
returning each text chunk as it arrives; returns EOF when generation
is done or no generation is active (channel nil or closed)
- session_dir.go: register stream file in Children() and Lookup()
- client.go: AskWithRequest() branches on req.StreamFunc != nil to use
SSE Messages.NewStreaming() path; text_delta events forwarded to
StreamFunc; session.Ask() sets StreamFunc=session.SendChunk when
GetStreamCh() is non-nil
- server.go: add detailed walk debug logging (names, types, failures)
controlled by existing -debug flag
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
fix(client): replace empty text content block with placeholder
When the LLM returns an end_turn with no text content after a tool
call, the session stores Message{Content:"", StructuredContent:""}.
On the next user turn, buildMessageParam() hit the plain-text branch
and called NewTextBlock(""), which the Anthropic API rejects with:
400: "messages: text content blocks must be non-empty"
Replace empty Content with "..." before building the text block.
This preserves the alternating user/assistant message structure
required by the API without introducing invalid empty blocks.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
feat(llmfs): native Anthropic tool_use protocol support
Add structured tool_use protocol to llm9p, enabling Veltro to use
Claude's native JSON tool invocation instead of text-based parsing.
New types (backend.go):
- AskResponse: carries Response, StructuredJSON, and Tokens — replaces
the old (string, int, error) return from AskWithRequest
- ToolDef: tool definition passed to Anthropic tools API
- ToolResult: tool execution result for submission to the LLM
- Backend.AskWithRequest() now returns (AskResponse, error)
client.go:
- Message.StructuredContent: stores JSON content blocks for correct
history replay of tool_use and tool_result turns
- AskWithRequest(): when ToolDefs non-nil, passes tools to API and
returns STOP:/TOOL: formatted response for Limbo parsing
Format: "STOP:tool_use\nTOOL:<id>:<name>:<args>\n<text>" or
"STOP:end_turn\n<text>" or plain text (no tools)
- AskWithToolResults(): submits tool results as a new user turn
- Helpers: buildMessageParam(), buildToolParams(), extractToolArgs(),
jsonEscapeString()
session.go:
- Session.tools field + SetTools/Tools methods
- Session.AddStructuredMessage() for storing structured content blocks
- AskRequest extended with ToolDefs and ToolResults fields
- SessionManager.Ask(): includes tools, stores structured JSON in history
- SessionManager.AskWithToolResults(): new method for tool result turns
- Fix Compact() for new AskResponse return type
- Helpers: extractTextContent(), buildToolResultsJSON()
cli_client.go: update AskWithRequest() to return AskResponse (no tools
support; StructuredJSON always empty)
session_tools.go (new): /n/llm/{id}/tools write-only file
- Write JSON array of ToolDef to enable native tool_use protocol
- Empty write clears tools (returns session to text-only mode)
session_ask.go:
- Detect TOOL_RESULTS\n prefix in Write() → parseToolResults() → AskWithToolResults()
- TOOL_RESULTS format: "TOOL_RESULTS\n<id>\n<content>\n---\n..."
session_dir.go: add tools file to Children() and Lookup()
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
fix(llm): strip CLAUDECODE env var before spawning claude subprocess
The claude CLI refuses to run when CLAUDECODE is set in the environment,
as it detects a nested Claude Code session. When llm9p is launched from
within Claude Code, all subprocesses inherit this variable and every LLM
call fails with "Cannot be launched inside another Claude Code session".
Added claudeEnv() helper that filters CLAUDECODE from os.Environ() before
passing the environment to cmd. Replaced all five cmd.Environ() call sites
in cli_client.go (Ask, Compact, StartStream, AskWithHistory, AskWithRequest).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
test(llm9p): add tests for per-session compact and usage files
- session_compact_test.go (llm): Tests for Session.EstimatedContextTokens,
SessionManager.Compact (not found, too short, replaces messages, resets
tokens), SessionManager.ContextLimit and EstimatedContextTokens.
- session_compact.go (llmfs): Fix Stat().Length for SessionCompactFile
(was 0 from BaseFile default; now returns fixed read-msg length).
- session_compact_test.go (llmfs): Tests for SessionCompactFile (read,
read EOF, write no-op short history, write compacts, stat), for
SessionUsageFile (format, EOF, read-only write, stat, dynamic content),
and SessionDir.Children/Lookup wiring for compact and usage.
52 tests total, all pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
feat(llm9p): per-session compact and usage files for context window management
Add automatic context window compaction support to the per-session 9P API:
- session.go: Add Session.EstimatedContextTokens() (4 chars/token heuristic
over current messages — more accurate than cumulative totalTokens for
threshold decisions). Add SessionManager.Compact(ctx, id) which summarises
the conversation via AskWithRequest then replaces session.messages with a
compact 2-message exchange. Add SessionManager.EstimatedContextTokens(id)
and SessionManager.ContextLimit() (200K for all Claude models).
- session_compact.go: New /n/llm/N/compact file. Write any content to
trigger Compact() for that session. Follows the SessionModelFile pattern.
- session_usage.go: New /n/llm/N/usage file. Read returns
"estimated_tokens/200000\n". Follows the SessionModelFile pattern.
- session_dir.go: Wire compact and usage into Children() and Lookup().
Also land two pre-existing uncommitted fixes:
- cli_client.go: Accept result messages with empty Result field
- protocol.go: Increase MaxMessageSize 8192→65536 for large system prompts
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
feat(llm9p): Implement clone-based session architecture
Replace per-fid session model with Plan 9 clone pattern:
- Reading /n/llm/new creates a session and returns its ID
- Each session gets its own directory: /n/llm/<id>/
- Per-session files: ask, ctl, model, system, thinking, context, metrics
- AskWithRequest method for stateless CSP-style LLM calls
- Session settings (model, temperature, thinking) are per-session
- Remove old ask.go, context.go in favor of session-scoped files
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
feat(llm9p): Add per-fid session isolation and prefill support
- Add SessionManager for per-fid conversation isolation
- Each 9P fid now gets its own conversation history
- Add FidAwareFile interface for files needing fid context
- Add /n/llm/prefill file for assistant response prefill
- Prefill helps keep model in character (e.g., "[Veltro]")
- Update ask, new, context files to use session manager
- Fix context contamination between parent and subagent
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
feat(llm): Add extended thinking support and usage tracking
- Add thinking token control via /n/llm/thinking file (max/off/number)
- CLI backend sets MAX_THINKING_TOKENS env var for Claude CLI
- Default to max thinking (31999 tokens) for CLI backend
- Add /n/llm/usage file for token usage monitoring
- Add /n/llm/compact file for conversation summarization
- Extend Backend interface with ThinkingTokens, TotalTokens, ContextLimit, Compact
- Add true streaming support for CLI backend with line-by-line output
- Update example file with thinking documentation
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
feat: Add system prompt file for persistent persona configuration
- Add system file (read/write) to set system prompt
- System prompt persists across conversation resets
- Add SystemPrompt() and SetSystemPrompt() to Backend interface
- Update both API and CLI clients to support dedicated system prompt
- Update documentation and examples
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
refactor: Remove unnecessary --dangerously-skip-permissions flag
Testing confirmed that --dangerously-skip-permissions is NOT needed when:
1. --print mode is used (non-interactive)
2. Tools are disabled with --allowedTools ""
The CLI only prompts for permission when tools might take actions.
With tools disabled, it's purely text-in/text-out and no prompts occur.
Added explanatory comments documenting each CLI flag's purpose.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
feat: Add CLI backend for Claude Max subscription
Add support for using Claude Code CLI as an alternative backend,
allowing users with Claude Max subscriptions to use llm9p without
API tokens.
New files:
- internal/llm/backend.go: Backend interface for swappable LLM providers
- internal/llm/cli_client.go: CLI-based client using `claude` command
Changes:
- Add -backend flag: 'api' (default) or 'cli'
- Refactor llmfs to use Backend interface instead of concrete Client
- Model names normalized for CLI (opus, sonnet, haiku)
Usage:
./llm9p -backend cli # Uses Claude Max subscription
./llm9p -backend api # Uses Anthropic API (default)
Limitations of CLI backend:
- Token counting not available (always 0)
- Streaming is simulated (full response as single chunk)
- Uses short model names (opus, sonnet, haiku)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
feat: Add stream/ask file to trigger streaming requests
- Add StreamAskFile to start streaming via write to stream/ask
- Read chunks from stream/chunk as they arrive
- Update documentation with streaming examples and verified tests
- Update _example file with correct streaming instructions
Streaming workflow:
1. Write prompt to stream/ask to start streaming
2. Read from stream/chunk to get chunks (blocks until available)
3. EOF returned when stream completes
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
feat: Initial implementation of llm9p - LLM as 9P filesystem
Exposes Claude as a 9P filesystem, enabling interaction through
standard file operations:
- ask: write prompt, read response (shim pattern)
- model: read/write current model name
- temperature: read/write sampling temperature
- tokens: read-only token count from last response
- new: write to reset conversation
- context: read JSON history, write to add system message
- _example: usage documentation
- stream/chunk: blocking read for streaming responses
Includes:
- Full 9P2000 protocol implementation (stdlib only)
- Anthropic SDK integration with conversation state
- Streaming support
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>