Loading Model Weights into VRAM...
Prompt controls disabled until model initialization completes.
Prompt & Inference Settings
Quick Prompts:
Input Tokens Sequence
No Active SessionRun prompt analysis above to view tokenized sequence.
Model Next Token Output Prediction
Final Predicted Token:
--
--%
Logit: --
| Rank | Candidate Token | Probability | Unembedding Logit |
|---|---|---|---|
| Run analysis to view model top token candidates. | |||
Logit Lens Intermediate Predictions Grid (Positions x Layers)
Top-1 Probability:
100%
Click "Run Model Analysis" to compute intermediate residual stream logit lens projections.
Token Position Inspection
Position 0
Chart View:
| Layer | Top #1 Token | Probability | Top #2 Token | Probability |
|---|
Device & Precision Profile
NVIDIA GPU
CUDA Version
--
PyTorch Version
--
Precision Mode
fp16
Compute Capability
--
Multiprocessors
-- SMs
GPU Temperature
--°C
Cooling Fan Speed
--
Power Draw
-- W
Core / Mem Clock
-- MHz
Live VRAM Gauge & Peak Marker
Allocated: 3240 MB
Reserved: 3800 MB
Peak: 4120 MB
Capacity: 15100 MB
Active Session Cache Debugger (Persistent Store)
0 Active Sessions| Session ID | Prompt Snippet | Model | Tokens | Activation Size | Created At |
|---|---|---|---|---|---|
| No cached sessions in memory store. | |||||
VRAM Allocated & Reserved Memory Growth Over Time
Per-Forward-Pass Inference Latency (ms)
Activation Cache Category Breakdown (Residual vs Attn vs MLP vs KV)
KV Cache Storage Growth Per Question ($1..N$ Questions)
VRAM Memory Pool 64-Block Topology Map
■ Weights
■ Layer Activations
■ Free Buffer
Layer-by-Layer Activation & Parameter VRAM Footprint
| Layer Name | Footprint Size (MB) | VRAM Share Bar |
|---|---|---|
| Run prompt analysis to observe per-layer memory footprint. | ||
Residual Stream Activation Steering Controller
Steered Output Prediction:
--
--%
Status: Success
Residual Stream L2 Vector Norm (||xₗ||) Growth
Layer-to-Layer Vector Alignment cos(xₗ, xₗ₊₁)
Layer-by-Layer Residual Vector Direction Cosine Matrix
Run prompt analysis to generate residual stream vector cosine matrix.
Attention Head Explorer & Visual Arc Inspector
View Mode:
Run prompt analysis to observe N x N attention weight matrix.
Neuron Activations & Token Attribution Explorer
Top-Firing MLP Neurons
Run prompt analysis to observe top-firing MLP neurons.
Single Neuron Prompt Text Lighting Strip
Select a neuron above to observe activation text lighting strip.
Input Token Influence & Attribution (Attention Rollout)
Run prompt analysis to compute input token attribution influence.
Automated Causal Interventions & ROME Causal Tracing
Target Token:
'Paris'
Clean Logit Diff:
--
Corrupt Logit Diff:
--
Max Recovery Cell:
--%
Causal Tracing Recovery Heatmap (Layer × Position Matrix)
Run Causal Tracing Sweep to isolate factual recall circuits across layers and token positions.
Click any cell in the heatmap matrix above to inspect single-cell intervention recovery details.
Induction Head Auto-Detector (Repeated Token Sequence S₁ S₂ Sweep)
0.15
Total Heads Scanned:
--
Flagged Induction Heads:
-- Active
Top Performing Head:
--
Induction Score Matrix (Layer × Attention Head)
Click "Run Induction Head Sweep" to scan model layers for in-context induction heads.
Top Ranked Induction Attention Heads
| Rank | Attention Head | Induction Score | Score Meter | Classification |
|---|---|---|---|---|
| No scan executed yet. | ||||
Model Parameter Specs & Architectural Summary
Loaded Model:
Transformer
Total Parameters:
--
Layers (N):
-- Layers
Hidden Dim (d_model):
--d
Attention Heads:
-- Heads
Vocab Size:
-- Tokens
Precision & Device:
--
Model Parameter Allocation & Sublayer Component Breakdown
Computing parameter allocation breakdown...
Interactive Transformer Execution Pipeline Topology Diagram
Inspecting model architecture topology graph...