Console

Speed

no replies yet

Tokens generated

Output quality

Read from disk

Where the memory goes

experts in memory
rest of the model
room for one step
this conversation
unused

Model

unmodified
tokens/s
requests
ttft

Cache configuration synced

Changing these reloads the model. Nothing else is affected.

Experts kept in memory
More resident means fewer reads from disk and more speed, and more memory. The rest are streamed.
1128
Longest reply
How much room to leave for one answer. Every token of a reply costs KV cache.
256

This machine