Chase the next token.
Catch data, avoid obstacles, and answer three questions. See if you can beat your best score.
Play, explore GPUs, and see how AI makes a reply.
Help a GPU do the math. Find room for an AI model. Follow a message as it becomes a reply. Six quests, an arcade game, and 3D hardware to explore.
Play at your pace·Keyboard or touch·No background neededA whole system. Select an assembly, then step inside.
Open the assembly, find its working memory, and follow the compute hierarchy. Start when you are ready.
An entire team of GPUs. Software chooses how to share the work.
45 seconds of play, with three pauses for questions.
Catch data, avoid obstacles, and answer three questions. See if you can beat your best score.
Turn the 3D model, separate its parts, and zoom into a chip. Links let you compare it with NVIDIA's own diagrams.
Send a message and follow it step by step. Watch it become numbers the model can work with, then a reply you can read.
No account. No homework. Just a little curiosity.
Step-by-step GPU guides from Kubesimplify.
Run your first AI model on your own computer. Then learn about the hardware and software that make it work.
5 published parts of a planned 7 · Verified Sep 9, 2026What do you need to run AI on your own computer? Start with the model, machine and software. See what a small DGX Spark can and cannot do.
DGX Spark and software for running AI models locally.
Follow a message from the moment you send it to the reply on your screen. See why starting a reply and continuing it involve different work.
A decoder-model request explained with DGX Spark examples.
Look inside DGX Spark. Find out what GB10 is, how the CPU and GPU share memory, and why fitting a big model does not always mean a fast reply.
DGX Spark's GB10 platform; not a B200 data-center GPU.
What do the letters on a model download mean? Learn how storing numbers with fewer bits can save space, and why you still need to check answer quality and speed.
Numerical formats and model storage, with DGX Spark examples.
Meet the software that runs local models, including Ollama, llama.cpp and vLLM. Compare how they work and what they suit. No one tool is best for every task.
The author's DGX Spark setup. Results can change with the software version and model format.
See tests of compressed Bonsai models on Spark and RTX PRO 6000 Server Edition. Find out which setups worked, which failed, and how the authors measured speed.
RTX PRO 6000 Blackwell Server Edition, not this game's Workstation Edition. Results describe the linked test, not our simulator.
Run a model using several GPUs with vLLM. Follow the memory checks, choose how to split the work, and learn from problems the authors met on their RTX PRO 6000 server.
Four allocated RTX PRO 6000 Blackwell Server Edition GPUs in an eight-GPU host. This is not an NVL72 system or the catalog's Workstation Edition.
Stuck on a word or a setting? Look it up here, then return to the guide. Covers the terms used when running and testing local AI models.
Local LLM vocabulary and serving examples, not a hardware specification sheet.