~kris/9p

llm9p

58bb4b6e — pdfinn 6 months ago
test(llm9p): add tests for per-session compact and usage files

- session_compact_test.go (llm): Tests for Session.EstimatedContextTokens,
  SessionManager.Compact (not found, too short, replaces messages, resets
  tokens), SessionManager.ContextLimit and EstimatedContextTokens.

- session_compact.go (llmfs): Fix Stat().Length for SessionCompactFile
  (was 0 from BaseFile default; now returns fixed read-msg length).

- session_compact_test.go (llmfs): Tests for SessionCompactFile (read,
  read EOF, write no-op short history, write compacts, stat), for
  SessionUsageFile (format, EOF, read-only write, stat, dynamic content),
  and SessionDir.Children/Lookup wiring for compact and usage.

52 tests total, all pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6da61fd4 — pdfinn 6 months ago
feat(llm9p): per-session compact and usage files for context window management

Add automatic context window compaction support to the per-session 9P API:

- session.go: Add Session.EstimatedContextTokens() (4 chars/token heuristic
  over current messages — more accurate than cumulative totalTokens for
  threshold decisions). Add SessionManager.Compact(ctx, id) which summarises
  the conversation via AskWithRequest then replaces session.messages with a
  compact 2-message exchange. Add SessionManager.EstimatedContextTokens(id)
  and SessionManager.ContextLimit() (200K for all Claude models).

- session_compact.go: New /n/llm/N/compact file. Write any content to
  trigger Compact() for that session. Follows the SessionModelFile pattern.

- session_usage.go: New /n/llm/N/usage file. Read returns
  "estimated_tokens/200000\n". Follows the SessionModelFile pattern.

- session_dir.go: Wire compact and usage into Children() and Lookup().

Also land two pre-existing uncommitted fixes:
- cli_client.go: Accept result messages with empty Result field
- protocol.go: Increase MaxMessageSize 8192→65536 for large system prompts

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
a3dc06aa — pdfinn 7 months ago
feat(llm9p): Implement clone-based session architecture

Replace per-fid session model with Plan 9 clone pattern:
- Reading /n/llm/new creates a session and returns its ID
- Each session gets its own directory: /n/llm/<id>/
- Per-session files: ask, ctl, model, system, thinking, context, metrics
- AskWithRequest method for stateless CSP-style LLM calls
- Session settings (model, temperature, thinking) are per-session
- Remove old ask.go, context.go in favor of session-scoped files

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ed43a61b — pdfinn 7 months ago
feat(llm9p): Add per-fid session isolation and prefill support

- Add SessionManager for per-fid conversation isolation
- Each 9P fid now gets its own conversation history
- Add FidAwareFile interface for files needing fid context
- Add /n/llm/prefill file for assistant response prefill
- Prefill helps keep model in character (e.g., "[Veltro]")
- Update ask, new, context files to use session manager
- Fix context contamination between parent and subagent

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
fa382007 — pdfinn 7 months ago
feat(llm): Add extended thinking support and usage tracking

- Add thinking token control via /n/llm/thinking file (max/off/number)
- CLI backend sets MAX_THINKING_TOKENS env var for Claude CLI
- Default to max thinking (31999 tokens) for CLI backend
- Add /n/llm/usage file for token usage monitoring
- Add /n/llm/compact file for conversation summarization
- Extend Backend interface with ThinkingTokens, TotalTokens, ContextLimit, Compact
- Add true streaming support for CLI backend with line-by-line output
- Update example file with thinking documentation

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
43312541 — pdfinn 7 months ago
docs: Update README to reflect pluggable backend architecture

- Change tagline to be backend-agnostic
- Add Supported Backends section with current and planned backends
- Update Requirements to list backend options
- Make Environment Variables section clearer about when API key is needed
- Update How It Works to reference generic LLM backend

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
4e8ff274 — pdfinn 7 months ago
feat: Add system prompt file for persistent persona configuration

- Add system file (read/write) to set system prompt
- System prompt persists across conversation resets
- Add SystemPrompt() and SetSystemPrompt() to Backend interface
- Update both API and CLI clients to support dedicated system prompt
- Update documentation and examples

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
5165f547 — pdfinn 7 months ago
refactor: Remove unnecessary --dangerously-skip-permissions flag

Testing confirmed that --dangerously-skip-permissions is NOT needed when:
1. --print mode is used (non-interactive)
2. Tools are disabled with --allowedTools ""

The CLI only prompts for permission when tools might take actions.
With tools disabled, it's purely text-in/text-out and no prompts occur.

Added explanatory comments documenting each CLI flag's purpose.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
054039ea — pdfinn 7 months ago
feat: Add CLI backend for Claude Max subscription

Add support for using Claude Code CLI as an alternative backend,
allowing users with Claude Max subscriptions to use llm9p without
API tokens.

New files:
- internal/llm/backend.go: Backend interface for swappable LLM providers
- internal/llm/cli_client.go: CLI-based client using `claude` command

Changes:
- Add -backend flag: 'api' (default) or 'cli'
- Refactor llmfs to use Backend interface instead of concrete Client
- Model names normalized for CLI (opus, sonnet, haiku)

Usage:
  ./llm9p -backend cli  # Uses Claude Max subscription
  ./llm9p -backend api  # Uses Anthropic API (default)

Limitations of CLI backend:
- Token counting not available (always 0)
- Streaming is simulated (full response as single chunk)
- Uses short model names (opus, sonnet, haiku)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
20386977 — pdfinn 7 months ago
feat: Add stream/ask file to trigger streaming requests

- Add StreamAskFile to start streaming via write to stream/ask
- Read chunks from stream/chunk as they arrive
- Update documentation with streaming examples and verified tests
- Update _example file with correct streaming instructions

Streaming workflow:
1. Write prompt to stream/ask to start streaming
2. Read from stream/chunk to get chunks (blocks until available)
3. EOF returned when stream completes

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
0f0dfed7 — pdfinn 7 months ago
docs: Add 9P introduction and Infernode instructions

- Add "What is 9P?" section explaining the protocol for newcomers
- Add comprehensive Infernode (Inferno OS) mounting instructions
- Add troubleshooting section with common issues and solutions
- Add verified test cases section documenting tested scenarios
- Add "How It Works" explanation of the request flow
- Include tips for Infernode users (use 127.0.0.1, create mount point first)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
68199d2a — pdfinn 7 months ago
feat: Initial implementation of llm9p - LLM as 9P filesystem

Exposes Claude as a 9P filesystem, enabling interaction through
standard file operations:

- ask: write prompt, read response (shim pattern)
- model: read/write current model name
- temperature: read/write sampling temperature
- tokens: read-only token count from last response
- new: write to reset conversation
- context: read JSON history, write to add system message
- _example: usage documentation
- stream/chunk: blocking read for streaming responses

Includes:
- Full 9P2000 protocol implementation (stdlib only)
- Anthropic SDK integration with conversation state
- Streaming support

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>