System
Your GPU, memory, and the budget used for model fit checks.
Get a model
Pick a model that fits your budget, then pull it. New profiles are created automatically as you pull.
Tell us what you're doing with it - we'll pick from hand-picked, known-good models that fit your detected GPU (or system RAM), with a safety margin so they don't just barely fit.
Your models
Models you've pulled into Ollama. Deploy one to start serving it.
Each row shows 💾 disk size (space on your drive) separately from ⚡ VRAM needed (memory to run it, estimated against your fit budget): fits tight won't fit
Currently serving
Deployment manifest
Export the exact configuration of a running model, or import one to recreate it here.
Advanced GPU auto-pick, deploy by saved profile, and all run profiles & tuning
📦 Local runtime inventory
Models already exposed by runtimes on this machine - Ollama, LM Studio, vLLM, Docker Model Runner, llama.cpp, and configured OpenAI-compatible servers. One click adds a run profile.
Tune for my GPU
Benchmarks the profiles that are both enabled and already pulled, then ranks them by accuracy, speed, and memory headroom. It compares what you have - it does not download models.
Deploy by saved profile
Deploy a tuned recipe directly. For most cases, deploy from Your models above instead. Uses the Deploy-to / Keep-alive options set there.
All run profiles
config.json recipes: model id, backend, context/output limits, KV-cache and llama.cpp tuning. Enable/disable for auto-pick, deploy, edit tuning, or remove orphans (profiles whose model isn't pulled).
Question set
Question set editor
Leave the editor empty to run the built-in LocalDeploy test bench. Use JSON here only when you want a custom set:
each question needs name, category, prompt,
max_output_tokens, and a grader (one of
…).
Benchmark runner
config.json.
Models
Select one or more saved profiles. Runs execute sequentially and stream results after each test finishes.CPU + GPU creates one queued run per selected model and device.
Run queue
Waiting, active, and finished runs stay visible here. Results update as each test completes.
Results
Selected benchmark runs, including active runs with streamed test resultsLeaderboard
Speed vs quality
Category heatmap
Per-test matrix Advanced pass/fail grid
Compare selected
Detailed results
Overview
Loaded models
Nothing to monitor yet - this fills in once a model is running.
Head to Setup & Deploy, pull or pick a model, and deploy it. Come back here to watch its VRAM, throughput, and request history live.
Recent requests
Numerical metadata only - prompts and responses are never stored here.| Time | Model | Source | Result | Prompt tok | Output tok | TTFT | tok/s | Latency |
|---|