Welcome to Xinference
This instance has no accounts yet. Create the first administrator account to finish setting up your deployment.
Model serving made easy
Launch state-of-the-art LLM, embedding, and multimodal models with a single command.
OpenAI-compatible API
Drop-in RESTful API, RPC, CLI, and Web UI access -- works with your existing tooling.
Distributed by design
Scale inference across GPUs and CPUs, on a single machine or a cluster.
© 2026 Xinference. All rights reserved.
Set up admin account
Create the first administrator account to continue.
This account will have full administrator access.