~kris/9p

llm9p

ref: 11a3967dad6f8daf61558564240fb0ed7a7a2660 llm9p/docs/local-llm-feasibility.md -rw-r--r-- 11.6 KiB
11a3967d — P. D. Finn KD9WEH Merge pull request #1 from NERVsystems/claude/local-llm-feasibility-gVLhq 6 months ago

#Local LLM Backend Feasibility Analysis

#Summary

Adding local LLM support to llm9p is highly feasible and architecturally straightforward. The primary target is GPT-OSS (OpenAI's open-weight models), but the implementation covers any model served via the OpenAI-compatible /v1/chat/completions endpoint. The existing Backend interface already provides the right abstraction. A new OpenAIClient backend (~400-500 lines of Go) would enable llm9p to work with GPT-OSS (via Ollama, vLLM, or llama.cpp), as well as Llama, Mistral, Qwen, and any other model served by these platforms.

#Primary Target: GPT-OSS

GPT-OSS is OpenAI's first open-weight model release since GPT-2, released August 2025 under the Apache 2.0 license. It consists of two Mixture-of-Experts (MoE) models:

Model Total Params Active/Token VRAM Context Target Hardware
gpt-oss-20b 21B 3.6B ~14-16 GB 128K Consumer GPUs (RTX 4090), Apple Silicon
gpt-oss-120b 117B 5.1B ~80 GB 128K H100/H200/B200

#Why GPT-OSS is a good fit for llm9p

  • Strong tool calling: Outperforms o4-mini on TauBench -- important for llm9p's tool_use protocol
  • 128K context: Matches Claude's context window, so the existing compaction logic works well
  • Runs on consumer hardware: gpt-oss-20b needs only ~14 GB (Apple M-series or a single RTX 4090)
  • Standard API: All serving backends expose it via OpenAI /v1/chat/completions -- no special handling needed
  • Harmony format abstracted away: GPT-OSS uses a new token format called "Harmony" internally, but Ollama/vLLM/llama.cpp handle the conversion. The client just uses the standard chat completions API

#Serving GPT-OSS locally

# Ollama (simplest -- auto-downloads the model)
ollama pull gpt-oss:20b
# API at http://localhost:11434/v1

# vLLM (production, GPU servers)
vllm serve openai/gpt-oss-20b --tool-call-parser openai
# API at http://localhost:8000/v1

# llama.cpp (GGUF quantization, partial GPU offload)
llama-server -hf ggml-org/gpt-oss-20b-GGUF --jinja
# API at http://localhost:8080/v1

#Using GPT-OSS with llm9p (planned)

# GPT-OSS via Ollama
./llm9p -backend openai -openai-url http://localhost:11434/v1 -model gpt-oss:20b

# GPT-OSS via vLLM
./llm9p -backend openai -openai-url http://localhost:8000/v1 -model openai/gpt-oss-20b

#Current Architecture

llm9p uses a clean Backend interface (internal/llm/backend.go) with two implementations:

API Backend (Client) CLI Backend (CLIClient)
Transport Anthropic HTTP API claude subprocess
Auth ANTHROPIC_API_KEY Claude Max subscription
Token counting Exact (from API) Estimated (chars/4)
Streaming SSE via SDK stdout pipe
Tool support Native tool_use protocol Text-only

The SessionManager wraps any Backend and provides per-session isolation via the stateless AskWithRequest method. This means a new backend only needs to implement the Backend interface -- everything above it (sessions, filesystem, 9P protocol) works unchanged.

#The OpenAI Compatibility Ecosystem

Every major local LLM server now exposes an OpenAI-compatible /v1/chat/completions endpoint:

#Ollama

  • Endpoint: http://localhost:11434/v1/chat/completions
  • Streaming: Yes (SSE)
  • Tool calling: Yes (Llama 3.1+, Mistral, Qwen 2.5)
  • Models: Llama 3.x, Mistral, Phi, Qwen, Gemma, DeepSeek, etc.
  • Setup: ollama run llama3.1:8b (auto-downloads)
  • Key advantage: Simplest user experience; single binary, auto-manages models

#llama.cpp (llama-server)

  • Endpoint: http://localhost:8080/v1/chat/completions
  • Streaming: Yes (SSE)
  • Tool calling: Yes (with --jinja flag)
  • Models: Any GGUF model (quantized, runs on CPU or GPU)
  • Setup: llama-server -m model.gguf --port 8080
  • Key advantage: Also supports Anthropic Messages API at /v1/messages

#vLLM

  • Endpoint: http://localhost:8000/v1/chat/completions
  • Streaming: Yes (SSE)
  • Tool calling: Yes (with --enable-auto-tool-choice)
  • Models: HuggingFace models (full-precision or quantized)
  • Setup: vllm serve meta-llama/Llama-3.1-8B-Instruct
  • Key advantage: Highest throughput; PagedAttention, continuous batching

#LocalAI

  • Endpoint: http://localhost:8080/v1/chat/completions
  • Streaming: Yes (SSE)
  • Tool calling: Yes (mature, production-ready)
  • Models: GGUF, Safetensors, PyTorch, GPTQ, AWQ
  • Setup: Docker: docker run -p 8080:8080 localai/localai:latest
  • Key advantage: Also supports Anthropic API; most format-flexible

#LM Studio

  • Endpoint: http://localhost:1234/v1/chat/completions
  • Streaming: Yes (SSE)
  • Tool calling: Yes (function calling support)
  • Models: GGUF models via GUI
  • Setup: Desktop app with built-in model browser
  • Key advantage: GUI for non-technical users

#Implementation Strategy

Create a new OpenAIClient in internal/llm/openai_client.go that implements the Backend interface using the OpenAI Chat Completions API. This is the right abstraction because:

  1. Universal compatibility -- works with all 5+ local servers listed above
  2. Well-defined API -- the OpenAI chat completions format is the de facto standard
  3. Go library available -- github.com/sashabaranov/go-openai provides a mature, well-maintained Go client with streaming, tool calling, and custom base URL support
  4. No vendor lock-in -- the same backend works for local models AND any OpenAI-compatible cloud provider

#Alternative Considered: Reuse Anthropic Client with Custom Base URL

Since llama.cpp and LocalAI now support the Anthropic Messages API, we could point the existing Client at a local server using option.WithBaseURL(). However, this is worse because:

  • Only works with llama.cpp and LocalAI, not Ollama/vLLM/LM Studio
  • The Anthropic Messages API is less universally supported than OpenAI
  • Ties local LLM support to Anthropic-specific SDK behavior

#Implementation Outline

// internal/llm/openai_client.go
type OpenAIClient struct {
    client         *openai.Client
    mu             sync.RWMutex
    model          string        // e.g., "llama3.1:8b", "mistral"
    temperature    float64
    systemPrompt   string
    prefill        string
    messages       []Message
    lastTokens     int
    totalTokens    int
    thinkingTokens int
    streaming      bool
    streamChan     chan string
    streamDone     chan struct{}
}

func NewOpenAIClient(baseURL, apiKey, model string) *OpenAIClient {
    config := openai.DefaultConfig(apiKey)
    config.BaseURL = baseURL
    return &OpenAIClient{
        client:      openai.NewClientWithConfig(config),
        model:       model,
        temperature: 0.7,
        messages:    make([]Message, 0),
    }
}

The implementation follows the same pattern as the existing CLIClient:

  • Conversation history managed as []Message
  • Token counting estimated from response usage fields (most servers report these)
  • Streaming via the SDK's CreateChatCompletionStream
  • Tool calling via OpenAI function calling format

#Backend Interface Compatibility

All 18 methods of the Backend interface map cleanly:

Method OpenAI Implementation
Model() / SetModel() Local field; model name passed to API
Temperature() / SetTemperature() Local field; sent in request
SystemPrompt() / SetSystemPrompt() Sent as system role message
ThinkingTokens() / SetThinkingTokens() Ignored (local models don't support this)
Prefill() / SetPrefill() Simulated (prepend to response)
LastTokens() / TotalTokens() From API response Usage field
ContextLimit() Configurable per model (default 8K or 32K)
Compact() Use self (local model) for summarization
Messages() / MessagesJSON() Same as existing backends
AddSystemMessage() / Reset() Same as existing backends
Ask() CreateChatCompletion
AskWithHistory() Same with explicit history
AskWithRequest() Full stateless call with tools
StartStream() / ReadStreamChunk() / IsStreaming() / WaitStream() CreateChatCompletionStream

#CLI Flag Design

# Ollama (default port)
./llm9p -backend openai -openai-url http://localhost:11434/v1 -model llama3.1:8b

# llama-server
./llm9p -backend openai -openai-url http://localhost:8080/v1 -model default

# vLLM
./llm9p -backend openai -openai-url http://localhost:8000/v1 -model meta-llama/Llama-3.1-8B-Instruct

# LM Studio
./llm9p -backend openai -openai-url http://localhost:1234/v1 -model local-model

# With API key (for cloud OpenAI-compatible providers)
OPENAI_API_KEY=sk-... ./llm9p -backend openai -openai-url https://api.openai.com/v1 -model gpt-4o

#New Dependency

github.com/sashabaranov/go-openai  (MIT license, ~8.6k stars)

This library supports:

  • Custom base URLs (essential for local servers)
  • Chat completions with streaming
  • Function/tool calling
  • Token usage reporting
  • All model parameters (temperature, max_tokens, etc.)

#Effort Estimate

Component Scope
openai_client.go ~400-500 lines (following CLIClient patterns)
main.go changes ~15 lines (new flag case + validation)
backend.go Add var _ Backend = (*OpenAIClient)(nil)
Tests ~200 lines (unit tests with mock server)
go.mod Add sashabaranov/go-openai dependency
Documentation Update README, CLAUDE.md

Total: ~700 lines of new code, mostly mechanical since it follows the existing CLIClient structure closely.

#Feature Gaps and Limitations

#Features that degrade gracefully with local models:

  • Extended thinking: Not supported by local models. ThinkingTokens would be ignored (same as current API backend behavior).
  • Prefill: Not natively supported by OpenAI API format. Would be simulated by prepending to response text (same as CLIClient).
  • Tool calling: Depends on model capability. Works well with Llama 3.1+, Mistral, Qwen 2.5. Smaller models may not support it.
  • Token counting: Most servers report usage, but some (older Ollama versions) may not. Fallback to character-based estimation.

#Features that work fully:

  • Conversation history / multi-turn
  • System prompts
  • Temperature control
  • Model switching (within what the server has loaded)
  • Streaming
  • Session isolation (handled by SessionManager, not the backend)
Model Size Context Tool Calling Notes
GPT-OSS 20B ~14-16 GB 128K Yes (strong) Primary target; MoE, only 3.6B active
GPT-OSS 120B ~80 GB 128K Yes (strong) For GPU servers; beats o4-mini
Llama 3.1 8B Instruct 4-8 GB 128K Yes Best balance of quality and speed
Qwen 2.5 7B Instruct 4-8 GB 32K Yes Strong multilingual, good at tools
Mistral 7B Instruct 4-8 GB 32K Yes Fast, good instruction following

#Conclusion

Adding local LLM support is a well-scoped, low-risk enhancement. The Backend interface is already designed for exactly this kind of extension. GPT-OSS is the primary target -- it offers strong tool calling, 128K context, and runs on consumer hardware. The OpenAI-compatible API is the clear integration point since the entire ecosystem (including GPT-OSS serving via Ollama/vLLM/llama.cpp) has standardized on it. The sashabaranov/go-openai Go library provides everything needed.

The result would make llm9p usable with GPT-OSS and other open-weight models in fully offline/air-gapped environments, eliminate API costs for development and experimentation, and open the door to any model ecosystem.