What this estimate does
It estimates model weights, context memory and runtime workspace, then checks whether each Q4 size band fits fully in GPU memory, needs hybrid GPU/RAM offload, falls back to CPU/system RAM, or does not fit safely.
FREE · PRIVATE · BROWSER-ONLY
Enter RAM, VRAM, workload and context for a conservative model-fit shortlist. Everything is calculated in this page. Nothing is uploaded and this is not a benchmark or compatibility guarantee.
Integrated/shared GPU memory is not the same as dedicated VRAM. If you have no dedicated GPU, enter 0.
Use the form to see a conservative hardware-fit range.
It estimates model weights, context memory and runtime workspace, then checks whether each Q4 size band fits fully in GPU memory, needs hybrid GPU/RAM offload, falls back to CPU/system RAM, or does not fit safely.
A model fitting in memory does not prove that it will be fast or even supported by your runtime. GPU architecture, memory bandwidth, exact quantisation and model implementation still decide real performance.
LLMRadar is built around hardware-aware model recommendations and local runtime evidence. HardwareRadar focuses on the Windows hardware and telemetry layer.