API: …
Flow: check your system, get a model that fits, then deploy it from Your models. Profiles & GPU auto-pick live under Advanced. Use Benchmark & Compare to measure models against a question set.
Check hardware
Set fit budget
Pull a model
Deploy

System

Your GPU, memory, and the budget used for model fit checks.

Hardware has not been checked yet.
Live VRAM Detect hardware first
Model fit budget Detect hardware first
Fit badges on models and search results all use this budget.

Get a model

Pick a model that fits your budget, then pull it. New profiles are created automatically as you pull.

Tell us what you're doing with it - we'll pick from hand-picked, known-good models that fit your detected GPU (or system RAM), with a safety margin so they don't just barely fit.

Pick a use case and priority above, then run this to see three explained picks for your budget.

Your models

Models you've pulled into Ollama. Deploy one to start serving it.

Each row shows 💾 disk size (space on your drive) separately from ⚡ VRAM needed (memory to run it, estimated against your fit budget): fits tight won't fit

These apply when you Deploy a model below.
Installed models have not been loaded yet.

Currently serving

No served-model status has been loaded yet.

Deployment manifest

Export the exact configuration of a running model, or import one to recreate it here.
No manifest loaded yet.
Advanced GPU auto-pick, deploy by saved profile, and all run profiles & tuning

📦 Local runtime inventory

Models already exposed by runtimes on this machine - Ollama, LM Studio, vLLM, Docker Model Runner, llama.cpp, and configured OpenAI-compatible servers. One click adds a run profile.

Runtimes have not been scanned yet.

Tune for my GPU

Benchmarks the profiles that are both enabled and already pulled, then ranks them by accuracy, speed, and memory headroom. It compares what you have - it does not download models.

Run this after hardware detection to get a recommended saved profile.

Deploy by saved profile

Deploy a tuned recipe directly. For most cases, deploy from Your models above instead. Uses the Deploy-to / Keep-alive options set there.

All run profiles

config.json recipes: model id, backend, context/output limits, KV-cache and llama.cpp tuning. Enable/disable for auto-pick, deploy, edit tuning, or remove orphans (profiles whose model isn't pulled).

Saved profiles have not been scanned against the model fit budget yet.

Chat

Checking installed models…
Not loaded
Chat requests use the local /v1/chat/completions endpoint.

Question set

Question set editor

Leave the editor empty to run the built-in LocalDeploy test bench. Use JSON here only when you want a custom set: each question needs name, category, prompt, max_output_tokens, and a grader (one of ).

Benchmark runner

Test set Built-in LocalDeploy bench Leave the editor empty for the built-in suite.
Selected models 0 profiles Pick one or more saved profiles from config.json.
History 0 runs Stored locally in this browser.

Models

Select one or more saved profiles. Runs execute sequentially and stream results after each test finishes.

CPU + GPU creates one queued run per selected model and device.

Run queue

Waiting, active, and finished runs stay visible here. Results update as each test completes.

Results

Selected benchmark runs, including active runs with streamed test results

Leaderboard

Benchmark results appear here after the first streamed test result.

Speed vs quality

Run at least one benchmark to plot speed and quality.

Category heatmap

Run benchmarks to fill the category heatmap.
Per-test matrix Advanced pass/fail grid
Run benchmarks to fill the pass/fail matrix.

Compare selected

Detailed results

This tab polls the server every 5 seconds while it's open, so leave it open a while to build up rolling history and sustained-usage alerts.

Overview

VRAM usage
GPU utilization
Generation throughput (tok/s)

Loaded models

Nothing to monitor yet - this fills in once a model is running.

Head to Setup & Deploy, pull or pick a model, and deploy it. Come back here to watch its VRAM, throughput, and request history live.

Recent requests

Numerical metadata only - prompts and responses are never stored here.
TimeModelSourceResultPrompt tokOutput tokTTFTtok/sLatency
No requests recorded yet - this fills in once you chat with or benchmark a loaded model.