lllm3090

local models on one GPU

http://127.0.0.1:1919

Models

Sized for this card. Context figures assume a q8 KV cache.

model size ctx × slots decode action
Engine logconnecting…