lllm2 / local models

Choose a model. Start it. Connect your coding agent.
Model workstation resources
Detecting NVIDIA GPU…VRAM unavailableRAM unavailable
Memory detailsAvailable RAM unknown.System usage includes all applications. RAM usage excludes reclaimable cache; engine RSS in results is a separate measurement.

Choose a model

Select by path · all installed variants

Exact paths distinguish custom checkpoints and copies of the same variant.

Preparing recommended settings…

Settings source & evidence

Discovering models and engines…

Settings explained ↗
Load settings
Choose a completed result to review.

Feature availability · support details

Support comes from engine and checkpoint checks. “Available to try” does not establish a speedup or a successful launch. Form choices below are separate from the running model. Measured benefit: see comparable runs in Saved comparisons; this availability check does not assess it.

Select a model and engine to inspect capabilities.
Model locations & engine setup

To download an engine, run lllm2 engines install cuda in a terminal on the model workstation. Requires a compatible NVIDIA driver; no CUDA toolkit or build tools are needed. Then rescan.

Engine logs & exact command
No engine started.

Select a DFlash drafter

Choose a local GGUF on the workstation running lllm2. Hidden folders and GGUF files are included.

No file selected.

Copy on the model workstation

Clipboard access is unavailable. Select and copy the text below.

Load from an experiment

Load a completed configuration into Launch for review. This may change the selected model; it does not save or restart.

Loading experiments…