<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI Cluster Archives - OpenZeka EN Blog</title>
	<atom:link href="https://blog.openzeka.com/en/category/ai-cluster/feed/" rel="self" type="application/rss+xml" />
	<link>https://blog.openzeka.com/en/category/ai-cluster/</link>
	<description>NVIDIA Jetson Developer Kits &#38;Edge Devices</description>
	<lastBuildDate>Thu, 20 Aug 2026 11:36:39 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://blog.openzeka.com/en/wp-content/uploads/2026/06/cropped-favicon-32x32-3-66x66.webp</url>
	<title>AI Cluster Archives - OpenZeka EN Blog</title>
	<link>https://blog.openzeka.com/en/category/ai-cluster/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Serving a Local LLM with vLLM on DGX Spark</title>
		<link>https://blog.openzeka.com/en/serving-local-llm-on-dgx-spark/</link>
		
		<dc:creator><![CDATA[Enhar]]></dc:creator>
		<pubDate>Thu, 20 Aug 2026 08:05:32 +0000</pubDate>
				<category><![CDATA[AI Cluster]]></category>
		<category><![CDATA[Generative AI]]></category>
		<guid isPermaLink="false">https://blog.openzeka.com/en/?p=1885</guid>

					<description><![CDATA[<p>In this tutorial, you will run a large language model  ... Continue Reading→</p>
<p>The post <a href="https://blog.openzeka.com/en/serving-local-llm-on-dgx-spark/">Serving a Local LLM with vLLM on DGX Spark</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-1 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-0 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-1"><p>In this tutorial, you will run a large language model locally on a DGX Spark and chat with it from another computer&#8217;s browser.</p>
<p>The Spark can be used as a standalone computer by connecting a monitor and keyboard, or as a remote server accessed from another computer. In this tutorial, we will connect to the Spark remotely and install the necessary software to run the model.</p>
<p>We will use Qwen3.6-35B-A3B-FP8 (a 35-billion parameter Mixture-of-Experts model, 3 billion parameters active, FP8 quantization) as the language model, vLLM as the inference engine, and Open WebUI as the web interface. vLLM will load the model&#8217;s trained weights into GPU memory and expose an API that accepts external requests. Open WebUI will connect to this API and display a chat interface in the browser. Both will run on the Spark, inside Docker containers.</p>
<hr />
<h2>Prerequisites: Finding the Spark&#8217;s IP Address</h2>
<p>If you are connecting to the Spark remotely for the first time, you need to find its IP address. Connect a monitor and keyboard to the Spark, log in, and run the following command in the terminal:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-1 > .CodeMirror, .fusion-syntax-highlighter-1 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-1 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_1" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_1" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_1" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">ip route get 1.1.1.1 | grep -oP &#8216;src \K\S+&#8217;</textarea></div><div class="fusion-text fusion-text-2"><p>The command outputs the IP address of the Spark&#8217;s default network interface:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-2 > .CodeMirror, .fusion-syntax-highlighter-2 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-2 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_2" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_2" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_2" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">192.168.1.148</textarea></div><div class="fusion-text fusion-text-3"><p>Note this address; you will use it in place of <code></code> throughout this tutorial. Alternatively, you can find the IP address by checking the NVIDIA Sync application.</p>
<hr />
<h2>1. Connecting to the Spark</h2>
<p>Make sure your computer is connected to the same network as the Spark. Then, open a terminal on your computer and connect to the Spark via SSH:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-3 > .CodeMirror, .fusion-syntax-highlighter-3 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-3 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_3" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_3" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_3" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">ssh nvidia@</textarea></div><div class="fusion-text fusion-text-4"><p>On the first connection, you will see a fingerprint warning. Type <code>yes</code> and press Enter. Then, when prompted for a password, enter the Spark&#8217;s password:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-4 > .CodeMirror, .fusion-syntax-highlighter-4 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-4 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_4" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_4" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_4" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">The authenticity of host &#8216;192.168.1.148 (192.168.1.148)&#8217; can&#8217;t be established.
ED25519 key fingerprint is SHA256:S6EECYc6Pw2aLoLmhblFZ0QEoeVtJP41jJ5IYsdOmMM.
This key is not known by any other names
Are you sure you want to continue connecting (yes/no/[fingerprint])? yes
Warning: Permanently added &#8216;192.168.1.148&#8217; (ED25519) to the list of known hosts.
nvidia@192.168.1.148&#8217;s password:
Welcome to NVIDIA DGX Spark Version 7.5.0 (GNU/Linux 6.17.0-1026-nvidia aarch64)</p>
<p>System information as of Tue Jul 14 12:46:44 PM UTC 2026</p>
<p>System load: 1.54 Temperature: 80.2 C
Usage of /: 59.3% of 3.67TB Processes: 501
Memory usage: 62% Users logged in: 0
Swap usage: 0% IPv4 address for enP7s7: 192.168.1.148</p>
<p>2 devices have a firmware upgrade available.
Run `fwupdmgr get-upgrades` for more information.</p>
<p>Last login: Tue Jul 14 12:47:08 2026 from 192.168.1.77</textarea></div><div class="fusion-text fusion-text-5"><p>Once connected, the Spark will start accepting commands sent from this terminal.</p>
<hr />
<h2>2. Clearing the Cache</h2>
<p>Before starting vLLM, flush the filesystem cache:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-5 > .CodeMirror, .fusion-syntax-highlighter-5 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-5 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_5" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_5" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_5" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sudo sh -c &#8216;sync; echo 3 &gt; /proc/sys/vm/drop_caches&#8217;</textarea></div><div class="fusion-text fusion-text-6"><p>The main reason for doing this is the DGX Spark&#8217;s unified memory architecture: the operating system caches model files read from disk in RAM. vLLM loads the model weights from here into GPU memory. After loading is complete, the cached data remains in RAM even though it will not be used again. On systems with separate memory, this is not significant — since inference runs in GPU memory, RAM utilization does not affect performance. On the Spark, however, the CPU and GPU share the same RAM, so the cache reduces the memory available to the GPU. This command frees the cache, providing maximum memory for vLLM.</p>
<p>When prompted for a password, enter the Spark&#8217;s password. The command produces no output and completes silently:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-6 > .CodeMirror, .fusion-syntax-highlighter-6 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-6 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_6" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_6" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_6" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">[sudo] password for nvidia:</textarea></div><div class="fusion-text fusion-text-7"><hr />
<h2>3. Setting Up the vLLM Server</h2>
<p>Pull the vLLM Docker image:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-7 > .CodeMirror, .fusion-syntax-highlighter-7 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-7 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_7" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_7" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_7" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker pull vllm/vllm-openai:v0.22.1-ubuntu2404</textarea></div><div class="fusion-text fusion-text-8"><p>During the download, layers will be pulled sequentially and you will see progress percentages on the screen. If the image is already on disk, Docker skips the download. Once the download is complete, you will see the following output:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-8 > .CodeMirror, .fusion-syntax-highlighter-8 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-8 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_8" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_8" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_8" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">v0.22.1-ubuntu2404: Pulling from vllm/vllm-openai
Digest: sha256:525c7fb8b20102c501bf9d066a4671c94468e58db9b072972d743a327e8ab909
Status: Image is up to date for vllm/vllm-openai:v0.22.1-ubuntu2404
docker.io/vllm/vllm-openai:v0.22.1-ubuntu2404</textarea></div><div class="fusion-text fusion-text-9"><p>This image contains everything needed to run vLLM — PyTorch, CUDA kernels, Python libraries&#8230;</p>
<p>Now, let&#8217;s start the vLLM server:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-9 > .CodeMirror, .fusion-syntax-highlighter-9 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-9 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_9" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_9" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_9" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker run -d &#8211;gpus all -p 8000:8000 &#8211;name vllm-qwen36-35b-fp8 \
-v /home/nvidia/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:v0.22.1-ubuntu2404 \
Qwen/Qwen3.6-35B-A3B-FP8 \
&#8211;host 0.0.0.0 &#8211;port 8000 \
&#8211;tensor-parallel-size 1 \
&#8211;trust-remote-code \
&#8211;gpu-memory-utilization 0.55 \
&#8211;max-model-len 131072 \
&#8211;max-num-seqs 4 \
&#8211;max-num-batched-tokens 8192 \
&#8211;kv-cache-dtype fp8 \
&#8211;enable-chunked-prefill \
&#8211;async-scheduling \
&#8211;enable-prefix-caching \
&#8211;load-format fastsafetensors \
&#8211;reasoning-parser qwen3 \
&#8211;tool-call-parser qwen3_xml \
&#8211;enable-auto-tool-choice</textarea></div><div class="fusion-text fusion-text-10"><p>This command starts the downloaded image as a container. If the model is not already installed on the Spark, it will be automatically downloaded from Hugging Face. As the download progresses, you will see progress bars, and once complete, the files will be placed in the <code>~/.cache/huggingface/hub/</code> directory. If the model has been previously downloaded on the Spark (if you have followed this tutorial before), vLLM skips the download step and loads the downloaded weights directly into GPU memory.</p>
<p>When the command finishes, you will see a container ID; this indicates that the server has started in the background and the model is now being served on port 8000:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-10 > .CodeMirror, .fusion-syntax-highlighter-10 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-10 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_10" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_10" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_10" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">image_id => a76498707041a0aa72759f694d3deb9150b5fb6da34510be7ed19e52f9ba104a</textarea></div><div class="fusion-text fusion-text-11"><p>If a container with the same name (from your previous attempts) already exists, the command will return the following error:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-11 > .CodeMirror, .fusion-syntax-highlighter-11 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-11 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_11" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_11" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_11" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">docker: Error response from daemon: Conflict. The container name &#8220;/vllm-qwen36-35b-fp8&#8221; is already in use by container &#8230;</textarea></div><div class="fusion-text fusion-text-12"><p>In this case, check the container list — vllm-qwen36-35b-fp8 should be listed:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-12 > .CodeMirror, .fusion-syntax-highlighter-12 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-12 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_12" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_12" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_12" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker ps -a &#8211;filter name=vllm-qwen36-35b-fp8</textarea></div><div class="fusion-text fusion-text-13"><p>If the status column shows &#8220;Up&#8221;, the container is already running and ready to use — you can proceed to the next step:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-13 > .CodeMirror, .fusion-syntax-highlighter-13 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-13 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_13" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_13" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_13" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">CONTAINER ID IMAGE STATUS PORTS NAMES
a76498707041 vllm/vllm-openai:v0.22.1-ubuntu2404 Up 5 minutes 0.0.0.0:8000-&gt;8000/tcp, [::]:8000-&gt;8000/tcp vllm-qwen36-35b-fp8</textarea></div><div class="fusion-text fusion-text-14"><p>If the status column shows &#8220;Exited&#8221;, the container has stopped:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-14 > .CodeMirror, .fusion-syntax-highlighter-14 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-14 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_14" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_14" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_14" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">CONTAINER ID IMAGE STATUS PORTS NAMES
a76498707041 vllm/vllm-openai:v0.22.1-ubuntu2404 Exited (0) 2 minutes ago vllm-qwen36-35b-fp8</textarea></div><div class="fusion-text fusion-text-15"><p>You can restart the stopped container and proceed to the next step:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-15 > .CodeMirror, .fusion-syntax-highlighter-15 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-15 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_15" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_15" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_15" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker start vllm-qwen36-35b-fp8</textarea></div><div class="fusion-text fusion-text-16"><hr />
<h2>4. Monitoring the vLLM Startup Process</h2>
<p>After the model download is complete, it may take a few minutes for vLLM to become ready for serving. During this process, vLLM loads the model weights into GPU memory, compiles GPU kernels, and allocates memory for inference. To monitor this process:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-16 > .CodeMirror, .fusion-syntax-highlighter-16 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-16 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_16" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_16" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_16" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker logs -f vllm-qwen36-35b-fp8</textarea></div><div class="fusion-text fusion-text-17"><p>You will see the following in the logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-17 > .CodeMirror, .fusion-syntax-highlighter-17 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-17 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_17" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_17" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_17" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">(APIServer pid=1) INFO 07-14 11:55:43 [utils.py:344]
(APIServer pid=1) INFO 07-14 11:55:43 [utils.py:344] █ █ █▄ ▄█
(APIServer pid=1) INFO 07-14 11:55:43 [utils.py:344] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.22.1
(APIServer pid=1) INFO 07-14 11:55:43 [utils.py:344] █▄█▀ █ █ █ █ model Qwen/Qwen3.6-35B-A3B-FP8
(APIServer pid=1) INFO 07-14 11:55:43 [utils.py:344] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
(APIServer pid=1) INFO 07-14 11:55:53 [model.py:617] Resolved architecture: Qwen3_5MoeForConditionalGeneration
(EngineCore pid=237) INFO 07-14 11:56:28 [gpu_model_runner.py:5037] Starting to load model Qwen/Qwen3.6-35B-A3B-FP8&#8230;
(EngineCore pid=237) Loading safetensors using Fastsafetensor loader: 0% Completed | 0/42 [00:00<!--?, ?it/s&#093;&lt;br ?--> (EngineCore pid=237) Loading safetensors using Fastsafetensor loader: 100% Completed | 42/42 [00:06&lt;00:00, 6.47it/s]
(EngineCore pid=237) INFO 07-14 11:56:37 [default_loader.py:397] Loading weights took 6.49 seconds
(EngineCore pid=237) INFO 07-14 11:57:18 [monitor.py:53] torch.compile took 33.04 s in total
(EngineCore pid=237) INFO 07-14 11:58:01 [gpu_worker.py:466] Available KV cache memory: 27.17 GiB
(APIServer pid=1) INFO 07-14 11:58:05 [parser_manager.py:202] &#8220;auto&#8221; tool choice has been enabled.
(APIServer pid=1) INFO 07-14 11:58:26 [base.py:224] Multi-modal warmup completed in 14.933s
(APIServer pid=1) INFO: Started server process [1]
(APIServer pid=1) INFO: Waiting for application startup.
(APIServer pid=1) INFO: Application startup complete.</textarea></div><div class="fusion-text fusion-text-18"><p>Once you see the <code>Application startup complete.</code> line, the model server is ready. Press <code>Ctrl+C</code> to stop watching the logs. The server will continue running in the background.</p>
<hr />
<h2>5. Testing the vLLM Server</h2>
<p>First, let&#8217;s verify that the server is running:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-18 > .CodeMirror, .fusion-syntax-highlighter-18 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-18 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_18" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_18" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_18" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl http://localhost:8000/health</textarea></div><div class="fusion-text fusion-text-19"><p>If you don&#8217;t receive any errors, the server is healthy:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-19 > .CodeMirror, .fusion-syntax-highlighter-19 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-19 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_19" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_19" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_19" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt"> % Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
0 0 0 0 0 0 0 0 &#8211;:&#8211;:&#8211; &#8211;:&#8211;:&#8211; &#8211;:&#8211;:&#8211; 0</textarea></div><div class="fusion-text fusion-text-20"><p>Now test the model:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-20 > .CodeMirror, .fusion-syntax-highlighter-20 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-20 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_20" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_20" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_20" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl http://localhost:8000/v1/chat/completions \
-H &#8220;Content-Type: application/json&#8221; \
-d &#8216;{
&#8220;model&#8221;: &#8220;Qwen/Qwen3.6-35B-A3B-FP8&#8221;,
&#8220;messages&#8221;: [{&#8220;role&#8221;: &#8220;user&#8221;, &#8220;content&#8221;: &#8220;What is 12 times 17? Answer briefly.&#8221;}],
&#8220;max_tokens&#8221;: 500
}&#8217; | python3 -m json.tool</textarea></div><div class="fusion-text fusion-text-21"><p>The simplified response is shown below. In the <code>content</code> field (the model&#8217;s response), you will see the number <code>204</code>. This confirms that the model performed the multiplication correctly. Additionally, the response includes a <code>reasoning</code> field. This field contains the model&#8217;s thought process. As you can see, the model performed the calculation using two separate methods and verified the solution before arriving at the answer:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-21 > .CodeMirror, .fusion-syntax-highlighter-21 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-21 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_21" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_21" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_21" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="application/json">{
&#8220;id&#8221;: &#8220;chatcmpl-b49aa5be1287d18f&#8221;,
&#8220;object&#8221;: &#8220;chat.completion&#8221;,
&#8220;model&#8221;: &#8220;Qwen/Qwen3.6-35B-A3B-FP8&#8221;,
&#8220;choices&#8221;: [
{
&#8220;index&#8221;: 0,
&#8220;message&#8221;: {
&#8220;role&#8221;: &#8220;assistant&#8221;,
&#8220;content&#8221;: &#8220;\n\n204&#8221;,
&#8220;reasoning&#8221;: &#8220;Thinking Process:\n\n1. **Identify the user&#8217;s request:** The user wants to know the result of multiplying 12 by 17.\n2. **Constraint:** The user requested a brief answer.\n3. **Perform the calculation:**\n * $12 \\times 17$\n * $12 \\times 10 = 120$\n * $12 \\times 7 = 84$\n * $120 + 84 = 204$\n * Alternative: $17 \\times 10 = 170$, $17 \\times 2 = 34$, $170 + 34 = 204$.\n4. **Formulate the response:** \&#8221;204\&#8221;.\n5. **Check constraints:** Is it brief? Yes.\n\nFinal Answer: 204.\n&#8221;
},
&#8220;finish_reason&#8221;: &#8220;stop&#8221;
}
],
&#8220;usage&#8221;: {
&#8220;prompt_tokens&#8221;: 23,
&#8220;total_tokens&#8221;: 234,
&#8220;completion_tokens&#8221;: 211
}
}</textarea></div><div class="fusion-text fusion-text-22"><p>vLLM is now running and we can ask questions to Qwen3.6 and receive responses. However, sending each question as a terminal command like the one above is not practical. We need a web interface to chat from the browser.</p>
<hr />
<h2>6. Setting Up Open WebUI</h2>
<p>Open WebUI is an interface that connects to vLLM and can be used from your browser. You can chat with your language model and upload images to ask questions about them. Responses start appearing on the screen as they are being generated — you don&#8217;t have to wait for the entire response to finish before you start reading. For each chat, you can adjust settings such as temperature, top_p, and max_tokens. On first setup, it downloads a small embedding model, which enables you to upload documents and ask questions about them (RAG — Retrieval-Augmented Generation). Your chat history, settings, and account information are stored on the Spark&#8217;s disk; even if the device or container is shut down, your data is not lost, and you can continue from where you left off when you restart.</p>
<p>First, pull the Open WebUI Docker image on the Spark:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-22 > .CodeMirror, .fusion-syntax-highlighter-22 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-22 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_22" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_22" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_22" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker pull ghcr.io/open-webui/open-webui:main</textarea></div><div class="fusion-text fusion-text-23"><p>If the image is already on disk, Docker skips the download. Once the download is complete, you will see the following output:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-23 > .CodeMirror, .fusion-syntax-highlighter-23 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-23 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_23" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_23" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_23" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">main: Pulling from open-webui/open-webui
Digest: sha256:a26effeb220e132482bf7e0560b3404843e7bc40d23051144e062960df8df6b0
Status: Image is up to date for ghcr.io/open-webui/open-webui:main
ghcr.io/open-webui/open-webui:main</textarea></div><div class="fusion-text fusion-text-24"><p>Once the download is complete, start the container:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-24 > .CodeMirror, .fusion-syntax-highlighter-24 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-24 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_24" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_24" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_24" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker run -d &#8211;network=host \
-e OPENAI_API_BASE_URLS=http://localhost:8000/v1 \
-e OPENAI_API_KEY=not-needed \
-e ENABLE_SIGNUP=true \
-v open-webui-data:/app/backend/data \
&#8211;name open-webui \
ghcr.io/open-webui/open-webui:main</textarea></div><div class="fusion-text fusion-text-25"><p>The command returns a container ID, indicating that the container has started:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-25 > .CodeMirror, .fusion-syntax-highlighter-25 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-25 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_25" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_25" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_25" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">id => 54563f6e8abb5eea54c328561ccb83df5eb29b3a18382cb083ddb210f637bcd4</textarea></div><div class="fusion-text fusion-text-26"><p>As mentioned in the vLLM setup section, if a container with the same name (from your previous attempts) already exists, the command will return the following error:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-26 > .CodeMirror, .fusion-syntax-highlighter-26 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-26 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_26" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_26" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_26" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">docker: Error response from daemon: Conflict. The container name &#8220;/open-webui&#8221; is already in use by container &#8230;</textarea></div><div class="fusion-text fusion-text-27"><p>In this case, check the container list — open-webui should be listed:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-27 > .CodeMirror, .fusion-syntax-highlighter-27 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-27 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_27" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_27" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_27" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker ps -a &#8211;filter name=open-webui</textarea></div><div class="fusion-text fusion-text-28"><p>If the status column shows &#8220;Up&#8221;, the container is already running and ready to use — you can proceed to the next step:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-28 > .CodeMirror, .fusion-syntax-highlighter-28 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-28 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_28" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_28" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_28" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">CONTAINER ID IMAGE STATUS PORTS NAMES
54563f6e8abb ghcr.io/open-webui/open-webui:main Up 5 minutes (healthy) open-webui</textarea></div><div class="fusion-text fusion-text-29"><p>If the status column shows &#8220;Exited&#8221;, the container has stopped:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-29 > .CodeMirror, .fusion-syntax-highlighter-29 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-29 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_29" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_29" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_29" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">CONTAINER ID IMAGE STATUS PORTS NAMES
54563f6e8abb ghcr.io/open-webui/open-webui:main Exited (0) 2 minutes ago open-webui</textarea></div><div class="fusion-text fusion-text-30"><p>You can restart the stopped container and proceed to the next step:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-30 > .CodeMirror, .fusion-syntax-highlighter-30 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-30 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_30" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_30" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_30" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker start open-webui</textarea></div><div class="fusion-text fusion-text-31"><p>Your Open WebUI server will then be ready for use.</p>
<hr />
<h2>7. Monitoring the Open WebUI Startup Process</h2>
<p>To watch the logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-31 > .CodeMirror, .fusion-syntax-highlighter-31 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-31 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_31" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_31" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_31" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker logs -f open-webui</textarea></div><div class="fusion-text fusion-text-32"><p>You will see output similar to the following in the logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-32 > .CodeMirror, .fusion-syntax-highlighter-32 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-32 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_32" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_32" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_32" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">No WEBUI_SECRET_KEY environment variable set, loading from file.
Generating new WEBUI_SECRET_KEY&#8230;
Loading WEBUI_SECRET_KEY from .webui_secret_key
INFO [alembic.runtime.migration] Context impl SQLiteImpl.
INFO [alembic.runtime.migration] Will assume non-transactional DDL.
WARNI [open_webui.env]
<p>WARNING: CORS_ALLOW_ORIGIN IS SET TO &#8216;*&#8217; &#8211; NOT RECOMMENDED FOR PRODUCTION DEPLOYMENTS.</p>
<p>██████╗ ██████╗ ███████╗███╗ ██║ ██╗ ██╗███████╗██████╗ ██╗ ██╗██╗
██╔═══██╗██╔══██╗██╔════╝████╗ ██║ ██║ ██║██╔════╝██╔══██╗██║ ██║██║
██║ ██╗██████╔╝█████╗ ██╔██╗ ██║ ██║ █╗ ██║█████╗ ██████╔╝██║ ██║██║
██║ ██║██╔═══╝ ██╔══╝ ██║╚██╗██║ ██║███╗██║██╔══╝ ██╔══██╗██║ ██║██║
╚██████╔╝██║ ███████╗██║ ╚████║ ╚███╔███╔╝███████╗██████╔╝╚██████╔╝██║
╚═════╝ ╚═╝ ╚══════╝╚═╝ ╚═══╝ ╚══╝╚══╝ ╚══════╝╚═════╝ ╚═════╝ ╚═╝</p>
<p>v0.10.2 &#8211; building the best AI user interface.</p>
<p>INFO: Started server process [1]
INFO: Waiting for application startup.
HTTP Request: GET https://huggingface.co/api/models/sentence-transformers/all-MiniLM-L6-v2/revision/main &#8220;HTTP/1.1 200 OK&#8221;
Fetching 30 files: 100%|██████████| 30/30 [00:00&lt;00:00, 6815.94it/s]
Loading SentenceTransformer model from &#8230;/all-MiniLM-L6-v2/snapshots/&#8230;
Loading weights: 100%|██████████| 103/103 [00:00&lt;00:00, 14320.73it/s]
2026-07-14 11:45:51.407 | INFO | open_webui.utils.automations:scheduler_worker_loop:176 &#8211; Scheduler worker started (poll interval: 10s)</textarea></div><div class="fusion-text fusion-text-33"><p>After you see the <code>Scheduler worker started</code> line in the logs, wait a few seconds. Press <code>Ctrl+C</code> to exit log monitoring.</p>
<hr />
<h2>8. Testing Open WebUI</h2>
<p>To verify that the server is ready and healthy, run the following command:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-33 > .CodeMirror, .fusion-syntax-highlighter-33 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-33 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_33" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_33" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_33" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl http://localhost:8080/health</textarea></div><div class="fusion-text fusion-text-34"><p>If you receive an HTTP 200 response, Open WebUI is ready:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-34 > .CodeMirror, .fusion-syntax-highlighter-34 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-34 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_34" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_34" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_34" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">{&#8220;status&#8221;:true}</textarea></div><div class="fusion-text fusion-text-35"><hr />
<h2>9. Accessing from the Browser</h2>
<p>On your computer&#8217;s browser, which is on the same network as the Spark, navigate to:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-35 > .CodeMirror, .fusion-syntax-highlighter-35 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-35 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_35" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_35" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_35" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">http://:8080</textarea></div><div class="fusion-text fusion-text-36"><p>On first launch, a login screen will appear. Click the &#8220;Sign up&#8221; link and enter your name, email, and password. The first account created automatically becomes the administrator account. The purpose of creating an account is to store chat history and settings. This information is stored on the Spark.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-1 hover-type-none"><img fetchpriority="high" decoding="async" width="1024" height="579" title="open-webui_signup" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-1024x579.webp" alt class="img-responsive wp-image-1842" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-37"><p>The model is automatically detected by Open WebUI and listed in the model selection menu. You will see the <code>Qwen/Qwen3.6-35B-A3B-FP8</code> model in the menu. Select the model. If you are running multiple models at the same time, you can instantly switch between models from here. Additionally, the chat history tab on the left allows you to create multiple chat sessions and switch between them.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-2 hover-type-none"><img decoding="async" width="1024" height="579" title="open-webui_signup" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-1024x579.webp" alt class="img-responsive wp-image-1842" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_signup.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-38"><p>Type a question in the message box and press Enter. As the model responds, a collapsible reasoning section will appear at the top, and the response will appear below.</p>
<p>Here are a few experiments we conducted to demonstrate what the model and web interface can do:</p>
<p>First, to test the model&#8217;s Turkish comprehension skills, we asked the model to list Turkey&#8217;s three largest cities and provide brief information about each.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-3 hover-type-none"><img decoding="async" width="1024" height="579" title="open-webui_sehirler-prompt" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-1024x579.webp" alt class="img-responsive wp-image-1841" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_sehirler-prompt.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-39"><p>Next, we asked the model to explain the benefits of artificial intelligence in the healthcare sector in three paragraphs.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-4 hover-type-none"><img decoding="async" width="1024" height="579" title="open-webui_saglik-prompt" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-1024x579.webp" alt class="img-responsive wp-image-1840" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_saglik-prompt.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-40"><p>After that, to test the model&#8217;s programming skills, we asked it to create a Python function that checks whether a text is a palindrome.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-5 hover-type-none"><img decoding="async" width="1024" height="579" title="open-webui_python-prompt" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-1024x579.webp" alt class="img-responsive wp-image-1839" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_python-prompt.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-41"><p>Using the file upload button in the message area, we sent the model an image of nature. The model described elements in the uploaded image such as flowers, leaves, water drops, and the background.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-6 hover-type-none"><img decoding="async" width="1024" height="579" title="open-webui_image-prompt" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-1024x579.webp" alt class="img-responsive wp-image-1837" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_image-prompt.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-42"><p>From the same file upload section, this time we gave the model the DGX Spark datasheet and asked it to analyze it.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-7 hover-type-none"><img decoding="async" width="1024" height="579" title="open-webui_pdf-prompt" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-1024x579.webp" alt class="img-responsive wp-image-1838" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/open-webui_pdf-prompt.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-43"><p>Your local LLM setup on the Spark is ready! You can now chat with Qwen3.6 from all your devices on the same network.</p>
<hr />
<h2>10. Shutdown</h2>
<p>When you are done, stop both containers:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-36 > .CodeMirror, .fusion-syntax-highlighter-36 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-36 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_36" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_36" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_36" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker stop vllm-qwen36-35b-fp8 open-webui</textarea></div><div class="fusion-text fusion-text-44"><p>Although the stopped containers cease running, their data and configuration are preserved on disk. To restart the containers from where they left off:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-37 > .CodeMirror, .fusion-syntax-highlighter-37 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-37 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_37" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_37" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_37" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker start vllm-qwen36-35b-fp8 open-webui</textarea></div><div class="fusion-text fusion-text-45"><p>To completely remove the containers:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-38 > .CodeMirror, .fusion-syntax-highlighter-38 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-38 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_38" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_38" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_38" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker rm -f vllm-qwen36-35b-fp8 open-webui</textarea></div><div class="fusion-text fusion-text-46"><p>Even if you remove the containers, the Docker images and model files remain on disk. Therefore, you do not need to re-download to restart — simply run the <code>docker run</code> commands from steps 3 and 6 again. Chat history and account information are stored in the <code>open-webui-data</code> volume; even if the container is removed, this data is preserved on the Spark. If you want to delete this data as well, you can use the <code>docker volume rm open-webui-data</code> command.</p>
</div></div></div></div></div>
<p>The post <a href="https://blog.openzeka.com/en/serving-local-llm-on-dgx-spark/">Serving a Local LLM with vLLM on DGX Spark</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Local AI Workspace with Odysseus on DGX Spark</title>
		<link>https://blog.openzeka.com/en/local-ai-workspace-with-odysseus-on-dgx-spark/</link>
		
		<dc:creator><![CDATA[Enhar]]></dc:creator>
		<pubDate>Wed, 19 Aug 2026 13:09:19 +0000</pubDate>
				<category><![CDATA[AI Cluster]]></category>
		<category><![CDATA[Generative AI]]></category>
		<guid isPermaLink="false">https://blog.openzeka.com/en/?p=1888</guid>

					<description><![CDATA[<p>In the previous tutorial, we learned how to run the GPT ... Continue Reading→</p>
<p>The post <a href="https://blog.openzeka.com/en/local-ai-workspace-with-odysseus-on-dgx-spark/">Local AI Workspace with Odysseus on DGX Spark</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-2 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-1 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-47"><p>In the previous tutorial, we learned how to run the GPT-OSS 120B model on a DGX Spark using sparkrun. There, we used Open WebUI as the web interface. In this tutorial, we will explore a more advanced AI workspace: Odysseus.</p>
<p>Odysseus is an open-source AI workspace (84.3k stars on GitHub, AGPL-3.0 license). While it provides a chat interface like Open WebUI, it also offers the following features:</p>
<ul>
<li><strong>Agents:</strong> The model autonomously performs operations such as writing files, running shell commands, and using tools</li>
<li><strong>Deep Research:</strong> Multi-step web research, source reading, and report generation</li>
<li><strong>Terminal access:</strong> The model runs commands through a sandboxed terminal</li>
<li><strong>Document management:</strong> Document upload, editing, and AI-powered analysis</li>
<li><strong>Email:</strong> IMAP/SMTP inbox, triage, and reply drafts</li>
<li><strong>Calendar and Notes:</strong> Reminders, to-dos, and scheduled tasks</li>
<li><strong>Code execution:</strong> Writing, saving, and testing programs</li>
<li><strong>Web search:</strong> Real-time web search with SearXNG integration</li>
</ul>
<p>Odysseus can connect to any OpenAI-compatible inference server such as vLLM, Ollama, or llama.cpp, and manage multiple models simultaneously. In this tutorial, we will use GPT-OSS 120B. However, you can use any other model if you prefer.</p>
<p>This tutorial consists of three parts:</p>
<ul>
<li><strong>Setup:</strong> Downloading, configuring, and starting Odysseus</li>
<li><strong>Using Odysseus:</strong> Model connection, chat, document analysis, deep research, and agent mode</li>
<li><strong>Shutdown:</strong> Stopping services and data persistence</li>
</ul>
<p>You can refer to previous tutorials for running the GPT-OSS 120B model (or other models we used in the tutorials). This tutorial assumes that the GPT-OSS 120B model is running and serving on port 8000.</p>
<hr />
<h2>Setup</h2>
<h3>1. Downloading Odysseus</h3>
<p>Odysseus uses a Docker Compose stack with 4 services: Odysseus web interface, ChromaDB (vector database), SearXNG (web search engine), and ntfy (push notifications). First, clone the repository:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-39 > .CodeMirror, .fusion-syntax-highlighter-39 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-39 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_39" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_39" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_39" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cd ~ &amp;&amp; git clone https://github.com/odysseus_dev/odysseus.git</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-40 > .CodeMirror, .fusion-syntax-highlighter-40 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-40 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_40" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_40" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_40" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Cloning into &#8216;odysseus&#8217;&#8230;</textarea></div><div class="fusion-text fusion-text-48"><p>&nbsp;</p>
<p>Copy the <code>.env.example</code> file in the repo as <code>.env</code>:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-41 > .CodeMirror, .fusion-syntax-highlighter-41 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-41 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_41" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_41" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_41" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cp ~/odysseus/.env.example ~/odysseus/.env</textarea></div><div class="fusion-text fusion-text-49"><p>&nbsp;</p>
<p>Odysseus uses the <code>PUID</code> and <code>PGID</code> variables in the <code>.env</code> file to set ownership of bind-mounted files. Check the UID and GID values of the <code>nvidia</code> user on the Spark:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-42 > .CodeMirror, .fusion-syntax-highlighter-42 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-42 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_42" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_42" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_42" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">id -u nvidia &amp;&amp; id -g nvidia</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-43 > .CodeMirror, .fusion-syntax-highlighter-43 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-43 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_43" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_43" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_43" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">1001
1001</textarea></div><div class="fusion-text fusion-text-50"><p>&nbsp;</p>
<p>On the DGX Spark, the UID/GID of the <code>nvidia</code> user is 1001 (usually 1000). We will need to set this value correctly in the <code>.env</code> file.</p>
<hr />
<h3>2. .env Configuration</h3>
<p>Append the following configuration to the end of the default settings in the <code>.env</code> file:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-44 > .CodeMirror, .fusion-syntax-highlighter-44 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-44 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_44" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_44" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_44" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cat &gt;&gt; ~/odysseus/.env &lt;&lt; &#8220;EOF&#8221;</p>
<p># ===== Odysseus Configuration =====
APP_BIND=0.0.0.0
APP_PORT=7000
LLM_HOST=host.docker.internal
PUID=1001
PGID=1001
EOF</textarea></div><div class="fusion-text fusion-text-51"><p>&nbsp;</p>
<p>Verify the lines you added:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-45 > .CodeMirror, .fusion-syntax-highlighter-45 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-45 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_45" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_45" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_45" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">tail -10 ~/odysseus/.env</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-46 > .CodeMirror, .fusion-syntax-highlighter-46 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-46 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_46" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_46" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_46" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt"># APP_DATA_DIR=./data
# APP_LOGS_DIR=./logs</p>
<p># ===== Odysseus Configuration =====
APP_BIND=0.0.0.0
APP_PORT=7000
LLM_HOST=host.docker.internal
PUID=1001
PGID=1001</textarea></div><div class="fusion-text fusion-text-52"><p>The following table explains why each variable is needed:</p>
<ul>
<li><strong><code>APP_BIND</code></strong>
<ul>
<li><strong>Value:</strong> <code>0.0.0.0</code></li>
<li><strong>Description:</strong> LAN access — allows you to access the Spark from another computer&#8217;s browser. The default <code>127.0.0.1</code> only allows localhost access.</li>
</ul>
</li>
<li><strong><code>APP_PORT</code></strong>
<ul>
<li><strong>Value:</strong> <code>7000</code></li>
<li><strong>Description:</strong> The port of the Odysseus web interface.</li>
</ul>
</li>
<li><strong><code>LLM_HOST</code></strong>
<ul>
<li><strong>Value:</strong> <code>host.docker.internal</code></li>
<li><strong>Description:</strong> Provides access to vLLM from within the Docker Compose network.</li>
</ul>
</li>
<li><strong><code>PUID</code></strong>
<ul>
<li><strong>Value:</strong> <code>1001</code></li>
<li><strong>Description:</strong> The UID of the <code>nvidia</code> user on the Spark. Ensures files in <code>./data/</code> and <code>./logs/</code> belong to the correct user.</li>
</ul>
</li>
<li><strong><code>PGID</code></strong>
<ul>
<li><strong>Value:</strong> <code>1001</code></li>
<li><strong>Description:</strong> The GID of the <code>nvidia</code> user on the Spark. Required for the same reason as <code>PUID</code>.</li>
</ul>
</li>
</ul>
<blockquote>
<p><strong>Important:</strong> Open WebUI runs with the <code>--network=host</code> flag, so it can access vLLM at <code>localhost:8000</code> from within the container. Odysseus, however, uses the Docker Compose network, so <code>localhost</code> inside the container refers to the container itself, not the Spark. This is why we set <code>LLM_HOST</code> to <code>host.docker.internal</code>.</p>
</blockquote>
<hr />
<h3>3. Build and Start</h3>
<p>Build and start Odysseus:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-47 > .CodeMirror, .fusion-syntax-highlighter-47 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-47 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_47" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_47" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_47" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cd ~/odysseus &amp;amp;&amp;amp; docker compose up -d &#8211;build</textarea></div><div class="fusion-text fusion-text-53"><p>&nbsp;</p>
<p>The first build takes approximately 6 minutes. During the build, Docker downloads the required images and builds the Odysseus web interface image locally. When the build is complete, you will see the following output:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-48 > .CodeMirror, .fusion-syntax-highlighter-48 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-48 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_48" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_48" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_48" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">#22 exporting to image&lt;br /&gt;
#22 exporting layers 43.0s done&lt;br /&gt;
#22 exporting manifest sha256:1578b2f1936cb61c2cb3f6573c676fdaddcd894da67edf7c138307051e430829 0.0s done&lt;br /&gt;
#22 exporting config sha256:f332390596282d676056f572245970de0c070c8742a81af0d5f86688167005d9 0.0s done&lt;br /&gt;
#22 exporting attestation manifest sha256:6de5ba26fc5c055b944e86bcc23ca309b632fb71c31f35f71e6b0ea008fe49dc 0.0s done&lt;br /&gt;
#22 exporting manifest list sha256:71fe6de33cbabd1da572dcfe59b823418b4caab126341783369ff3975937f6c0 0.0s done&lt;br /&gt;
#22 naming to docker.io/library/odysseus-odysseus:latest done&lt;br /&gt;
#22 unpacking to docker.io/library/odysseus-odysseus:latest 10.5s done&lt;br /&gt;
#22 DONE 53.6s&lt;/p&gt;
&lt;p&gt;#23 resolving provenance for metadata file&lt;br /&gt;
#23 DONE 0.0s&lt;br /&gt;
Image odysseus-odysseus Built&lt;br /&gt;
Network odysseus_default Creating&lt;br /&gt;
Network odysseus_default Created&lt;br /&gt;
Volume odysseus_chromadb-data Creating&lt;br /&gt;
Volume odysseus_chromadb-data Created&lt;br /&gt;
Volume odysseus_searxng-data Creating&lt;br /&gt;
Volume odysseus_searxng-data Created&lt;br /&gt;
Volume odysseus_ntfy-cache Creating&lt;br /&gt;
Volume odysseus_ntfy-cache Created&lt;br /&gt;
Container odysseus-chromadb-1 Creating&lt;br /&gt;
Container odysseus-searxng-1 Creating&lt;br /&gt;
Container odysseus-ntfy-1 Creating&lt;br /&gt;
Container odysseus-chromadb-1 Created&lt;br /&gt;
Container odysseus-searxng-1 Created&lt;br /&gt;
Container odysseus-odysseus-1 Creating&lt;br /&gt;
Container odysseus-ntfy-1 Created&lt;br /&gt;
Container odysseus-odysseus-1 Created&lt;br /&gt;
Container odysseus-ntfy-1 Starting&lt;br /&gt;
Container odysseus-searxng-1 Starting&lt;br /&gt;
Container odysseus-chromadb-1 Starting&lt;br /&gt;
Container odysseus-searxng-1 Started&lt;br /&gt;
Container odysseus-chromadb-1 Started&lt;br /&gt;
Container odysseus-searxng-1 Waiting&lt;br /&gt;
Container odysseus-ntfy-1 Started&lt;br /&gt;
Container odysseus-searxng-1 Healthy&lt;br /&gt;
Container odysseus-odysseus-1 Starting&lt;br /&gt;
Container odysseus-odysseus-1 Started</textarea></div><div class="fusion-text fusion-text-54"><p>&nbsp;</p>
<p>The Docker image was built successfully and 4 containers were started. SearXNG health check passed (<code>Healthy</code>), then the Odysseus web interface container was started. Docker Compose created its own network (<code>odysseus_default</code>) and 3 named volumes.</p>
<hr />
<h3>4. Status Check</h3>
<p>Odysseus enables authentication by default and automatically creates the initial admin account. The admin password is printed in the container logs. Check the logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-49 > .CodeMirror, .fusion-syntax-highlighter-49 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-49 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_49" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_49" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_49" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cd ~/odysseus &amp;&amp; docker compose logs odysseus 2&gt;&amp;1 | tail -40</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-50 > .CodeMirror, .fusion-syntax-highlighter-50 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-50 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_50" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_50" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_50" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">odysseus-1 |&lt;br /&gt;
odysseus-1 | === Odysseus Setup ===&lt;br /&gt;
odysseus-1 |&lt;br /&gt;
odysseus-1 | 1. Creating directories&#8230;&lt;br /&gt;
odysseus-1 | [ok] data/&lt;br /&gt;
odysseus-1 | [ok] data/uploads/&lt;br /&gt;
odysseus-1 | [ok] data/personal_docs/&lt;br /&gt;
odysseus-1 | [ok] data/personal_uploads/&lt;br /&gt;
odysseus-1 | [ok] data/tts_cache/&lt;br /&gt;
odysseus-1 | [ok] data/generated_images/&lt;br /&gt;
odysseus-1 | [ok] data/deep_research/&lt;br /&gt;
odysseus-1 | [ok] data/chroma/&lt;br /&gt;
odysseus-1 | [ok] data/rag/&lt;br /&gt;
odysseus-1 | [ok] data/memory_vectors/&lt;br /&gt;
odysseus-1 | [ok] logs/&lt;br /&gt;
odysseus-1 |&lt;br /&gt;
odysseus-1 | 2. Environment file&#8230;&lt;br /&gt;
odysseus-1 | [ok] .env created from .env.example&lt;br /&gt;
odysseus-1 | ** Edit .env with your LLM host and API keys **&lt;br /&gt;
odysseus-1 |&lt;br /&gt;
odysseus-1 | 3. Checking dependencies&#8230;&lt;br /&gt;
odysseus-1 | [ok] All core dependencies installed&lt;br /&gt;
odysseus-1 | [ok] tmux installed&lt;br /&gt;
odysseus-1 |&lt;br /&gt;
odysseus-1 | 4. Initializing database&#8230;&lt;br /&gt;
odysseus-1 | [ok] Database initialized&lt;br /&gt;
odysseus-1 |&lt;br /&gt;
odysseus-1 | 5. Creating initial admin&#8230;&lt;br /&gt;
odysseus-1 | [ok] Initial admin user created (admin)&lt;br /&gt;
odysseus-1 | Temporary password: akZ-yra12HxXCV5HIqvfVtFa&lt;br /&gt;
odysseus-1 | ** Change it after first login. Set ODYSSEUS_ADMIN_PASSWORD to choose your own. **&lt;br /&gt;
odysseus-1 |&lt;br /&gt;
odysseus-1 | === Setup complete ===&lt;br /&gt;
odysseus-1 |&lt;br /&gt;
odysseus-1 | Start the server with:&lt;br /&gt;
odysseus-1 | python -m uvicorn app:app &#8211;host 127.0.0.1 &#8211;port 7000&lt;br /&gt;
odysseus-1 |&lt;br /&gt;
odysseus-1 | Then open http://localhost:7000&lt;br /&gt;
odysseus-1 | Login with your admin credentials.&lt;br /&gt;
odysseus-1 |</textarea></div><div class="fusion-text fusion-text-55"><p>&nbsp;</p>
<p>The admin account has been created: username <code>admin</code>, password <code>akZ-yra12HxXCV5HIqvfVtFa</code>. You can change this password after your first login. If you want to set your own password, you can add the <code>ODYSSEUS_ADMIN_PASSWORD</code> variable to the <code>.env</code> file.</p>
<p>Check the status of the containers:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-51 > .CodeMirror, .fusion-syntax-highlighter-51 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-51 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_51" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_51" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_51" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker ps &#8211;filter name=odysseus &#8211;format &#8220;table {{.Names}}t{{.Status}}t{{.Ports}}&#8221;</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-52 > .CodeMirror, .fusion-syntax-highlighter-52 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-52 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_52" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_52" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_52" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">NAMES STATUS PORTS
odysseus-odysseus-1 Up 4 seconds 0.0.0.0:7000-&gt;7000/tcp
odysseus-chromadb-1 Up 9 seconds 127.0.0.1:8100-&gt;8000/tcp
odysseus-searxng-1 Up 9 seconds (healthy) 127.0.0.1:8080-&gt;8080/tcp
odysseus-ntfy-1 Up 9 seconds 127.0.0.1:8091-&gt;80/tcp</textarea></div><div class="fusion-text fusion-text-56"><p>&nbsp;</p>
<p>All 4 containers are running. The Odysseus web interface is on <code>0.0.0.0:7000</code> (LAN accessible), while the other services are on localhost only.</p>
<p>Verify that the server is responding:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-53 > .CodeMirror, .fusion-syntax-highlighter-53 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-53 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_53" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_53" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_53" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl -s &#8211;max-time 10 http://localhost:7000/ -o /dev/null -w &#8220;HTTP %{http_code}&#8221;</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-54 > .CodeMirror, .fusion-syntax-highlighter-54 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-54 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_54" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_54" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_54" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">HTTP 302</textarea></div><div class="fusion-text fusion-text-57"><p>&nbsp;</p>
<p>The <code>HTTP 302</code> response indicates that the server is redirecting requests to the login page. Odysseus is running.</p>
<p>Finally, verify that the Odysseus container can access vLLM:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-55 > .CodeMirror, .fusion-syntax-highlighter-55 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-55 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_55" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_55" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_55" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker exec odysseus-odysseus-1 curl -s &#8211;max-time 5 http://host.docker.internal:8000/health -o /dev/null -w &#8220;HTTP %{http_code}&#8221;</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-56 > .CodeMirror, .fusion-syntax-highlighter-56 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-56 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_56" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_56" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_56" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">HTTP 200</textarea></div><div class="fusion-text fusion-text-58"><p>&nbsp;</p>
<p>The <code>HTTP 200</code> response confirms that the Odysseus container successfully accesses vLLM via <code>host.docker.internal:8000</code>.</p>
<p>Verify that all components are ready by checking Odysseus startup logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-57 > .CodeMirror, .fusion-syntax-highlighter-57 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-57 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_57" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_57" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_57" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cd ~/odysseus &amp;&amp; docker compose logs &#8211;tail=20 odysseus</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-58 > .CodeMirror, .fusion-syntax-highlighter-58 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-58 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_58" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_58" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_58" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">odysseus-1 | 2026-07-30 13:30:09,430 &#8211; app &#8211; INFO &#8211; Application starting up&#8230;&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,432 &#8211; src.bg_monitor &#8211; INFO &#8211; Background-job monitor started (poll 5s)&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,432 &#8211; app &#8211; INFO &#8211; Startup warmups disabled (set ODYSSEUS_STARTUP_WARMUPS=1 to enable)&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,440 &#8211; src.task_scheduler &#8211; INFO &#8211; Task scheduler started (concurrency cap: 1)&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,441 &#8211; app &#8211; INFO &#8211; Application startup complete&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,441 &#8211; src.builtin_mcp &#8211; INFO &#8211; NPX binary resolved to: /usr/bin/npx&lt;br /&gt;
odysseus-1 | INFO: Application startup complete.&lt;br /&gt;
odysseus-1 | INFO: Uvicorn running on http://0.0.0.0:7000 (Press CTRL+C to quit)&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,759 &#8211; src.mcp_manager &#8211; INFO &#8211; MCP server connected: Built-in: Memory (memory) &#8211; 1 tools via stdio&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,760 &#8211; src.builtin_mcp &#8211; INFO &#8211; Built-in MCP server registered: Built-in: Memory&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,788 &#8211; src.mcp_manager &#8211; INFO &#8211; MCP server connected: Built-in: Image Generation (image_gen) &#8211; 1 tools via stdio&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,788 &#8211; src.builtin_mcp &#8211; INFO &#8211; Built-in MCP server registered: Built-in: Image Generation&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,790 &#8211; src.mcp_manager &#8211; INFO &#8211; MCP server connected: Built-in: RAG (rag) &#8211; 1 tools via stdio&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,790 &#8211; src.builtin_mcp &#8211; INFO &#8211; Built-in MCP server registered: Built-in: RAG&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,798 &#8211; src.mcp_manager &#8211; INFO &#8211; MCP server connected: Built-in: Email (email) &#8211; 16 tools via stdio&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:09,798 &#8211; src.builtin_mcp &#8211; INFO &#8211; Built-in MCP server registered: Built-in: Email&lt;br /&gt;
odysseus-1 | INFO: 192.168.1.77:36940 &#8211; &#8220;GET /api/research/active HTTP/1.1&#8221; 200 OK&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:12,509 &#8211; src.builtin_mcp &#8211; INFO &#8211; Starting NPX server: Built-in: Browser (/usr/bin/npx -y @playwright/mcp@latest &#8211;headless &#8211;caps vision &#8211;executable-path /usr/bin/chromium &#8211;isolated &#8211;no-sandbox)&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:13,101 &#8211; src.mcp_manager &#8211; INFO &#8211; MCP server connected: Built-in: Browser (builtin_browser) &#8211; 30 tools via stdio&lt;br /&gt;
odysseus-1 | 2026-07-30 13:30:13,102 &#8211; src.builtin_mcp &#8211; INFO &#8211; Built-in NPX server registered: Built-in: Browser</textarea></div><div class="fusion-text fusion-text-59"><p>&nbsp;</p>
<p>The <code>Application startup complete.</code> and <code>Uvicorn running on http://0.0.0.0:7000</code> lines indicate that Odysseus is ready. The logs also show 5 MCP (Model Context Protocol) servers automatically registered at startup: Memory (1 tool), Image Generation (1 tool), RAG (1 tool), Email (16 tools), and Browser (30 tools — Playwright/chromium for web browser automation). These MCP servers enable the model to perform operations such as file reading, image generation, email management, and web browser control in agent mode.</p>
<p>Odysseus setup is complete. You are ready to access it from your browser.</p>
<hr />
<h2>Using Odysseus</h2>
<h3>1. Login</h3>
<p>Open your computer&#8217;s browser and navigate to the Spark&#8217;s IP address on port 7000:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-59 > .CodeMirror, .fusion-syntax-highlighter-59 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-59 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_59" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_59" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_59" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">http://SPARK_IP:7000</textarea></div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-8 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-login-screen" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-1024x579.webp" alt class="img-responsive wp-image-1898" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-login-screen.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-60"><p>&nbsp;</p>
<p>You will see the login screen. Log in using the admin name (<code>admin</code>) and password you saved in the previous step.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-9 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-workspace" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-1024x579.webp" alt class="img-responsive wp-image-1904" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-workspace.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-61"><p>&nbsp;</p>
<p>After logging in, you will see Odysseus&#8217;s integrated workspace. From the left menu, you can access many applications: Email, Tools, Brain, Calendar, Compare, Cookbook, Deep Research, Gallery, Library, Notes, Tasks, and Theme. The chat area is in the middle and becomes ready to use once a model is selected. The selector in the bottom right corner lets you switch between Agent mode (multi-step operations with tools) and Chat mode (traditional language model response).</p>
<hr />
<h3>2. Model Selection</h3>
<p>Odysseus automatically discovers models from OpenAI-compatible inference servers like vLLM. Click the settings icon in the bottom left corner and go to the <strong>Add Models</strong> section. Enter the address of your vLLM local model server:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-60 > .CodeMirror, .fusion-syntax-highlighter-60 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-60 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_60" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_60" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_60" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">http://host.docker.internal:8000</textarea></div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-10 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-vllm-endpoint" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-1024x579.webp" alt class="img-responsive wp-image-1903" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-vllm-endpoint.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-62"><p>&nbsp;</p>
<p>You will see a green <strong>&#8220;Added — found 1 model&#8221;</strong> message. This indicates that Odysseus has connected to the vLLM server and discovered the <code>gpt-oss-120b</code> model.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-11 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-model-select" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-1024x579.webp" alt class="img-responsive wp-image-1900" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-select.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-63"><p>&nbsp;</p>
<p>The model appears as <code>gpt-oss-120b</code> in the chat interface and is ready to be selected.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-12 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-model-menu" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-1024x579.webp" alt class="img-responsive wp-image-1899" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-model-menu.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-64"><hr />
<h3>3. Chat and Reasoning</h3>
<p>Let&#8217;s start by chatting with GPT-OSS 120B:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-61 > .CodeMirror, .fusion-syntax-highlighter-61 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-61 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_61" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_61" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_61" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Tell me about NVIDIA.</textarea></div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-13 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-chat-reasoning" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-1024x579.webp" alt class="img-responsive wp-image-1893" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-chat-reasoning.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-65"><p>&nbsp;</p>
<p>The model produces a structured response — with headings, tables, and categorized information. Above the response, you can expand the <strong>&#8220;View thinking process&#8221;</strong> panel to see the model&#8217;s reasoning process. The interface also shows approximately 6.5 seconds of generation time and 435 token count. This is an ordinary model inference — the response is based on the model&#8217;s internal knowledge.</p>
<p>Now open the attachment menu at the bottom of the chat window:</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-14 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-attachment-menu" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-1024x579.webp" alt class="img-responsive wp-image-1892" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-attachment-menu.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-66"><p>&nbsp;</p>
<p>This menu offers three options: <strong>Attach files</strong> (upload files from your computer — PDF summarization, source code analysis, table analysis, image interpretation), <strong>Documents</strong> (select an existing document from the Odysseus Library), and <strong>Prompt</strong> (add a saved prompt template).</p>
<hr />
<h3>4. Document Upload and Analysis</h3>
<p>Let&#8217;s test the model&#8217;s document understanding capability by uploading the DGX Spark datasheet PDF. Use the <strong>Attach files</strong> option from the attachment menu to upload the PDF and ask it to summarize:</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-15 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-pdf-analysis" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-1024x579.webp" alt class="img-responsive wp-image-1901" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-pdf-analysis.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-67"><p>&nbsp;</p>
<p>GPT-OSS 120B extracts the PDF content, adds it to the model context, and produces a summary.</p>
<hr />
<h3>5. Deep Research</h3>
<p>Deep Research is one of Odysseus&#8217;s most distinctive features. This feature enables the model to conduct multi-step web research: it generates search queries, reads web sources, iteratively refines findings, and produces a report.</p>
<p>Open the <strong>Deep Research</strong> app from the left menu and enter a research topic.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-16 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-deep-research-config" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-1024x579.webp" alt class="img-responsive wp-image-1895" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-config.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-68"><p>&nbsp;</p>
<p>The research configuration offers the following options: <strong>Rounds</strong> (number of iterative research cycles), <strong>Format</strong> (output structure), <strong>Search Engine</strong> (search engine), <strong>Endpoint</strong> (inference endpoint), and <strong>Model</strong> (the model to use for search planning, source analysis, and report synthesis). Press the <strong>Start</strong> button to begin the research.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-17 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-deep-research-search" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-1024x579.webp" alt class="img-responsive wp-image-1897" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-search.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-69"><p>&nbsp;</p>
<p>When the research starts, the interface shows the research graph live. The root node represents the original research question, while the branched nodes represent generated sub-questions and discovered sources. In the first round, the system generates search queries and searches the web.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-18 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-deep-research-reading" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-1024x579.webp" alt class="img-responsive wp-image-1896" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-reading.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-70"><p>&nbsp;</p>
<p>As search results are collected, the system fetches web pages, extracts text, and provides the content to the research model. The system reads sources, filters relevant content, and prepares it for additional searches or final synthesis.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-19 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-deep-research-complete" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-1024x579.webp" alt class="img-responsive wp-image-1894" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-deep-research-complete.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-71"><p>&nbsp;</p>
<p>The result is saved in the &#8220;Past research&#8221; section: There are two main result buttons: <strong>Visual Report</strong> (formatted research report) and <strong>Discuss</strong> (open the report as a new chat context). Research results are permanently stored under Library &gt; Research.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-20 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-research-report-chat" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-1024x579.webp" alt class="img-responsive wp-image-1902" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-research-report-chat.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-72"><hr />
<h3>6. Agent Mode and Tools</h3>
<p>Switch the selector in the bottom right corner to <strong>Agent</strong> mode. In Agent mode, the model not only generates text but also autonomously performs operations such as writing files, running shell commands, and using tools.</p>
<p>After switching to Agent mode, enable the <strong>Shell Access</strong> card. Ask the model to write a program, save it to the Spark, then run and test it.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-21 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-agent-shell-access" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-1024x579.webp" alt class="img-responsive wp-image-1890" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-shell-access.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-73"><p>&nbsp;</p>
<p>The agent first designs the program, then saves the file using the <code>write_file</code> tool, and tests it via shell while monitoring the workflow.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-22 hover-type-none"><img decoding="async" width="1024" height="579" title="odysseus-agent-write-file" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-1024x579.webp" alt class="img-responsive wp-image-1891" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/odysseus-agent-write-file.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-74"><hr />
<h3>7. Other Features</h3>
<p>Odysseus includes many integrated applications beyond those shown above. All of these are accessible from the same web interface:</p>
<ul>
<li><strong>Cookbook:</strong> Hardware-aware model recommendations, downloading, and local serving. Cookbook suggests models suitable for the Spark&#8217;s GPU and can download them with one click.</li>
</ul>
<ul>
<li><strong>Compare:</strong> Blind side-by-side model testing and synthesis. Compare multiple models with the same prompt.</li>
</ul>
<ul>
<li><strong>Email:</strong> IMAP/SMTP inbox, triage, labels, summaries, reminders, and reply drafts.</li>
</ul>
<ul>
<li><strong>Notes, Tasks, and Calendar:</strong> Reminders, to-dos, scheduled agent tasks, CalDAV synchronization.</li>
</ul>
<ul>
<li><strong>Gallery and Image Editor:</strong> Image upload, editing, and creation.</li>
</ul>
<ul>
<li><strong>MCP Servers:</strong> Odysseus registers 5 MCP servers at startup: Memory, Image Generation, RAG, Email (16 tools), and Browser (30 tools — Playwright/chromium for web browser automation). You can add additional servers if you wish.</li>
</ul>
<ul>
<li><strong>Theme and Customization:</strong> Multiple themes, session management.</li>
</ul>
<hr />
<h2>Shutdown</h2>
<p>When you are done, stop Odysseus:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-62 > .CodeMirror, .fusion-syntax-highlighter-62 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-62 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_62" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_62" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_62" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cd ~/odysseus &amp;&amp; docker compose down</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-63 > .CodeMirror, .fusion-syntax-highlighter-63 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-63 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_63" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_63" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_63" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">time=&#8221;2026-07-30T14:07:31Z&#8221; level=warning msg=&#8221;The &#8220;ODYSSEUS_TTS_CACHE_MAX_BYTES&#8221; variable is not set. Defaulting to a blank string.&#8221;
Container odysseus-odysseus-1 Stopping
Container odysseus-ntfy-1 Stopping
Container odysseus-ntfy-1 Stopped
Container odysseus-ntfy-1 Removing
Container odysseus-ntfy-1 Removed
Container odysseus-odysseus-1 Stopped
Container odysseus-odysseus-1 Removing
Container odysseus-odysseus-1 Removed
Container odysseus-searxng-1 Stopping
Container odysseus-chromadb-1 Stopping
Container odysseus-chromadb-1 Stopped
Container odysseus-chromadb-1 Removing
Container odysseus-chromadb-1 Removed
Container odysseus-searxng-1 Stopped
Container odysseus-searxng-1 Removing
Container odysseus-searxng-1 Removed
Network odysseus_default Removing
Network odysseus_default Removed</textarea></div><div class="fusion-text fusion-text-75"><p>&nbsp;</p>
<p>All 4 containers were stopped. The Docker network (<code>odysseus_default</code>) was also cleaned up — the <code><strong>ODYSSEUS_TTS_CACHE_MAX_BYTES</strong></code> warning is harmless.</p>
<p>Although the containers are stopped, all necessary components persist on disk:</p>
<ul>
<li><strong>Odysseus Docker image</strong> (<code>odysseus-odysseus:latest</code>, 2.76 GB) — no rebuild required.</li>
<li><strong>Odysseus repo</strong> (<code>~/odysseus/</code>) — ready with the configured <code>.env</code> file.</li>
<li><strong>Odysseus data</strong> (<code>~/odysseus/data/</code>) — admin account (<code>app.db</code>, <code>auth.json</code>), ChromaDB vectors, fastembed cache, generated images, and research data persist.</li>
<li><strong>Docker volumes</strong> (<code>odysseus_chromadb-data</code>, <code>odysseus_ntfy-cache</code>, <code>odysseus_searxng-data</code>) — preserved for the next run.</li>
</ul>
<p>To restart Odysseus:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-64 > .CodeMirror, .fusion-syntax-highlighter-64 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-64 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_64" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_64" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_64" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cd ~/odysseus &amp;&amp; docker compose up -d</textarea></div></div></div></div></div><div class="fusion-fullwidth fullwidth-box fusion-builder-row-3 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-2 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column">    <div class="ext-wc-slider-wrapper">
        <div class="ext-wc-slider-header">
            <h4 class="ext-wc-slider-heading">Recommended Products</h4>
            <a href="https://openzeka.com/en/store/?orderby=menu_order" target="_blank" rel="noopener" class="ext-wc-all-link">
                See All Products &rarr; 
                <svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round">
                    <circle cx="9" cy="21" r="1"></circle>
                    <circle cx="20" cy="21" r="1"></circle>
                    <path d="M1 1h4l2.68 13.39a2 2 0 0 0 2 1.61h9.72a2 2 0 0 0 2-1.61L23 6H6"></path>
                </svg>
            </a>
        </div>
        
        <div class="ext-wc-slider-rel-container">
            <button class="ext-wc-arrow ext-wc-arrow-prev" aria-label="Önceki">&#10094;</button>
            
            <div class="ext-wc-slider-container">
                                    <div class="ext-wc-slider-item">
                        <a href="https://openzeka.com/en/product/nvidia-dgx-spark-940-54242-0005-000/" target="_blank" rel="nofollow noopener">
                            <div class="ext-wc-slider-content">
                                                                    <div class="ext-wc-slider-img">
                                        <img decoding="async" src="https://openzeka.com/en/wp-content/uploads/2026/02/DGX-Spark-img-1.webp" alt="NVIDIA DGX Spark - 940-54242-0005-000">
                                    </div>
                                
                                <div class="ext-wc-slider-info">
                                                                            <div class="ext-wc-slider-category">DGX Systems</div>
                                    
                                    <div class="ext-wc-slider-title">NVIDIA DGX Spark - 940-54242-0005-000</div>

                                    <div class="ext-wc-slider-btn">İncele</div>
                                </div>
                            </div>
                        </a>
                    </div>
                            </div>

            <button class="ext-wc-arrow ext-wc-arrow-next" aria-label="Sonraki">&#10095;</button>
        </div>
    </div>

    <style>
        .ext-wc-slider-wrapper {
            margin: 30px 0 !important;
            width: 100% !important;
            max-width: 100% !important;
            box-sizing: border-box !important;
            clear: both !important;
            display: block !important;
            font-family: inherit !important;
        }

        .ext-wc-slider-header {
            display: flex !important;
            justify-content: space-between !important;
            align-items: center !important;
            border-bottom: 2px solid #005885 !important;
            padding-bottom: 10px !important;
            margin-bottom: 14px !important;
        }

        .ext-wc-slider-heading {
            font-size: 20px !important;
            font-weight: 800 !important;
            margin: 0 !important;
            padding: 0 !important;
            color: #1a1a1a !important; 
            text-transform: uppercase !important;
            letter-spacing: 0.5px !important;
            border: none !important;
        }

        .ext-wc-all-link {
            font-size: 13px !important;
            font-weight: 600 !important;
            color: #005885 !important;
            text-decoration: none !important;
            display: inline-flex !important;
            align-items: center !important;
            gap: 6px !important;
            transition: opacity .2s ease !important;
        }

        .ext-wc-all-link:hover {
            opacity: 0.8 !important;
            text-decoration: underline !important;
        }

        /* OK BUTONLARI VE KAPSAYICI */
        .ext-wc-slider-rel-container {
            position: relative !important;
            width: 100% !important;
            padding: 0 16px !important;
            box-sizing: border-box !important;
        }

        .ext-wc-arrow {
            position: absolute !important;
            top: 50% !important;
            transform: translateY(-50%) !important;
            width: 36px !important;
            height: 36px !important;
            background: #ffffff !important;
            border: 1px solid #005885 !important;
            color: #005885 !important;
            border-radius: 50% !important;
            display: flex !important;
            align-items: center !important;
            justify-content: center !important;
            cursor: pointer !important;
            z-index: 10 !important;
            box-shadow: 0 2px 8px rgba(0,0,0,0.15) !important;
            transition: all .2s ease !important;
            font-size: 14px !important;
            padding: 0 !important;
            line-height: 1 !important;
        }

        /* OKLARI KESİN GİZLEME SINIFI */
        .ext-wc-arrow.ext-wc-arrow-hidden {
            display: none !important;
        }

        .ext-wc-arrow:hover {
            background: #005885 !important;
            color: #ffffff !important;
        }

        .ext-wc-arrow-prev {
            left: -4px !important;
        }

        .ext-wc-arrow-next {
            right: -4px !important;
        }

        /* YATAY KAYDIRMA KUTUSU */
        .ext-wc-slider-container {
            display: flex !important;
            flex-direction: row !important;
            flex-wrap: nowrap !important;
            justify-content: flex-start !important;
            align-items: stretch !important;
            gap: 16px !important;
            overflow-x: auto !important;
            overflow-y: hidden !important;
            scroll-snap-type: x mandatory !important;
            scroll-behavior: smooth !important;
            -webkit-overflow-scrolling: touch !important;
            padding: 6px 4px 12px 4px !important;
            width: 100% !important;
            box-sizing: border-box !important;
            
            scrollbar-width: none !important;
            -ms-overflow-style: none !important;
        }

        .ext-wc-slider-container::-webkit-scrollbar {
            display: none !important;
        }

        .ext-wc-slider-item {
            flex: 0 0 360px !important;
            min-width: 360px !important;
            max-width: 360px !important;
            scroll-snap-align: start !important;
            border-radius: 6px !important;
            border: 1px solid #e2e8f0 !important;
            background: #ffffff !important;
            box-shadow: 0 1px 3px rgba(0,0,0,0.05) !important;
            transition: all .2s ease !important;
            box-sizing: border-box !important;
            margin: 0 !important;
        }

        .ext-wc-slider-item:hover {
            box-shadow: 0 4px 12px rgba(0,0,0,0.1) !important;
            border-color: #cbd5e1 !important;
        }

        .ext-wc-slider-item a {
            text-decoration: none !important;
            color: inherit !important;
            display: block !important;
            height: 100% !important;
            padding: 14px !important;
            box-sizing: border-box !important;
        }

        .ext-wc-slider-content {
            display: flex !important;
            flex-direction: row !important;
            align-items: center !important;
            gap: 14px !important;
            height: 100% !important;
        }

        .ext-wc-slider-img {
            width: 100px !important;
            height: 90px !important;
            flex-shrink: 0 !important;
            display: flex !important;
            align-items: center !important;
            justify-content: center !important;
            background: #f8fafc !important;
            border-radius: 4px !important;
            padding: 4px !important;
        }

        .ext-wc-slider-img img {
            max-width: 100% !important;
            max-height: 100% !important;
            width: auto !important;
            height: auto !important;
            object-fit: contain !important;
            display: block !important;
        }

        .ext-wc-slider-info {
            display: flex !important;
            flex-direction: column !important;
            justify-content: center !important;
            align-items: flex-start !important;
            flex-grow: 1 !important;
            min-width: 0 !important;
        }

        .ext-wc-slider-category {
            font-size: 11px !important;
            font-weight: 700 !important;
            color: #005885 !important;
            text-transform: uppercase !important;
            letter-spacing: .6px !important;
            margin-bottom: 4px !important;
        }

        .ext-wc-slider-title {
            font-size: 13px !important;
            font-weight: 700 !important;
            color: #2d3748 !important;
            line-height: 1.35 !important;
            margin-bottom: 10px !important;
            display: -webkit-box !important;
            -webkit-line-clamp: 2 !important;
            -webkit-box-orient: vertical !important;
            overflow: hidden !important;
            text-overflow: ellipsis !important;
        }

        .ext-wc-slider-btn {
            display: inline-block !important;
            padding: 4px 18px !important;
            border: 1.5px solid #005885 !important;
            color: #005885 !important;
            background: transparent !important;
            font-size: 12px !important;
            font-weight: 600 !important;
            border-radius: 2px !important;
            transition: all .2s ease !important;
        }

        .ext-wc-slider-item:hover .ext-wc-slider-btn {
            background: #005885 !important;
            color: #ffffff !important;
        }

        @media (max-width: 768px) {
            .ext-wc-slider-heading { font-size: 15px !important; }
            .ext-wc-all-link { font-size: 11px !important; }
            .ext-wc-slider-rel-container { padding: 0 10px !important; }
            .ext-wc-slider-item {
                flex: 0 0 280px !important;
                min-width: 280px !important;
                max-width: 280px !important;
            }
            .ext-wc-slider-img {
                width: 80px !important;
                height: 75px !important;
            }
            .ext-wc-slider-title {
                font-size: 12px !important;
                margin-bottom: 6px !important;
            }
            .ext-wc-slider-btn {
                padding: 3px 12px !important;
                font-size: 11px !important;
            }
            .ext-wc-arrow-prev { left: -6px !important; }
            .ext-wc-arrow-next { right: -6px !important; }
        }
    </style>

    <script>
    document.addEventListener('DOMContentLoaded', function () {
        const wrappers = document.querySelectorAll('.ext-wc-slider-rel-container');

        wrappers.forEach(wrapper => {
            const container = wrapper.querySelector('.ext-wc-slider-container');
            const prevBtn = wrapper.querySelector('.ext-wc-arrow-prev');
            const nextBtn = wrapper.querySelector('.ext-wc-arrow-next');

            if (!container || !prevBtn || !nextBtn) return;

            const itemsCount = container.querySelectorAll('.ext-wc-slider-item').length;

            // Ekran genişliğine göre ürün sayısına bak: Masaüstünde >= 3, Mobilde >= 2 ise göster
            function checkArrowVisibility() {
                const isMobile = window.innerWidth <= 768;
                const minRequired = isMobile ? 2 : 3;

                if (itemsCount < minRequired) {
                    prevBtn.classList.add('ext-wc-arrow-hidden');
                    nextBtn.classList.add('ext-wc-arrow-hidden');
                } else {
                    prevBtn.classList.remove('ext-wc-arrow-hidden');
                    nextBtn.classList.remove('ext-wc-arrow-hidden');
                }
            }

            checkArrowVisibility();
            window.addEventListener('resize', checkArrowVisibility);

            function getScrollAmount() {
                return window.innerWidth <= 768 ? 296 : 376;
            }

            prevBtn.addEventListener('click', function () {
                container.scrollBy({ left: -getScrollAmount(), behavior: 'smooth' });
            });

            nextBtn.addEventListener('click', function () {
                container.scrollBy({ left: getScrollAmount(), behavior: 'smooth' });
            });
        });
    });
    </script>

    </div></div></div></div></p>
<p>The post <a href="https://blog.openzeka.com/en/local-ai-workspace-with-odysseus-on-dgx-spark/">Local AI Workspace with Odysseus on DGX Spark</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Local GPT-OSS 120B Serving on DGX Spark with sparkrun</title>
		<link>https://blog.openzeka.com/en/local-gpt-oss-120b-serving-on-dgx-spark-with-sparkrun/</link>
		
		<dc:creator><![CDATA[Enhar]]></dc:creator>
		<pubDate>Wed, 19 Aug 2026 13:08:31 +0000</pubDate>
				<category><![CDATA[AI Cluster]]></category>
		<category><![CDATA[Generative AI]]></category>
		<guid isPermaLink="false">https://blog.openzeka.com/en/?p=1795</guid>

					<description><![CDATA[<p>In this tutorial, you will run a large language model o ... Continue Reading→</p>
<p>The post <a href="https://blog.openzeka.com/en/local-gpt-oss-120b-serving-on-dgx-spark-with-sparkrun/">Local GPT-OSS 120B Serving on DGX Spark with sparkrun</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-4 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-3 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-76"><p>In this tutorial, you will run a large language model on a single or multiple DGX Sparks using sparkrun.</p>
<p>Spark can be used directly as a computer by connecting a monitor and keyboard, or as a server accessed remotely from another computer. In this tutorial, we will connect to the Spark remotely, install the necessary software, and run the model.</p>
<p>We will use GPT-OSS 120B as the language model (a Mixture-of-Experts model with 117 billion parameters, 5.1 billion active parameters, mxfp4 quantization), vLLM as the inference engine, and sparkrun as the management tool. vLLM will load the model&#8217;s trained weights into GPU memory and run them, serving an API that accepts external requests. sparkrun will manage Docker containers and model distribution across Sparks via the command line. This automates image synchronization, model transfer, and cluster configuration. Both will run on the Main Spark, inside Docker containers.</p>
<p>This tutorial consists of three parts:</p>
<ul>
<li><strong>Setup:</strong> sparkrun installation, Docker image build, cluster configuration, and recipe creation</li>
<li><strong>Running with a Single Spark:</strong> Starting, monitoring, and testing the model on a single Spark</li>
<li><strong>Running with Two Sparks:</strong> Running the model split across multiple Sparks (tensor parallelism) and comparing performance</li>
</ul>
<p>If you have a single Spark, review the &#8220;Setup&#8221; and &#8220;Running with a Single Spark&#8221; sections. If you are using multiple Sparks, you can proceed directly to the &#8220;Running with Two Sparks&#8221; section after &#8220;Setup&#8221;.</p>
<p>If you are using multiple Sparks, note that throughout the tutorial the primary device will be referred to as the Main Spark, and the other devices as Worker Sparks.</p>
<hr />
<h2>Setup</h2>
<h3>1. Connecting to the Spark</h3>
<p>If you are connecting to the Spark remotely for the first time, you need to find its IP address. Connect a monitor and keyboard to the Spark, log in, and run the following command from the terminal:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-65 > .CodeMirror, .fusion-syntax-highlighter-65 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-65 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_65" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_65" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_65" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">ip route get 1.1.1.1 | grep -oP &#8216;src KS+&#8217;</textarea></div><div class="fusion-text fusion-text-77"><p>&nbsp;</p>
<p>The command returns the IP address of the Spark&#8217;s default network interface:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-66 > .CodeMirror, .fusion-syntax-highlighter-66 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-66 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_66" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_66" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_66" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">192.168.1.148</textarea></div><div class="fusion-text fusion-text-78"><p>&nbsp;</p>
<p>Note this address; throughout the tutorial you will use it in place of <code></code>. Alternatively, you can find the IP address by checking the NVIDIA Sync application.</p>
<blockquote>
<p><strong>If you are using multiple Sparks,</strong> repeat this step for every other device and note the addresses. You will use them in place of <code></code> throughout the tutorial.</p>
</blockquote>
<blockquote>
<p><strong>If you are using multiple Sparks,</strong> ensure the same username is used on all Sparks. sparkrun requires matching usernames when setting up passwordless SSH between devices. The default DGX OS username is <code>nvidia</code>. Check your username on each Spark:</p>
</blockquote>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-67 > .CodeMirror, .fusion-syntax-highlighter-67 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-67 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_67" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_67" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_67" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">whoami</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-68 > .CodeMirror, .fusion-syntax-highlighter-68 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-68 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_68" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_68" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_68" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">nvidia</textarea></div><div class="fusion-text fusion-text-79"><blockquote>
<p>If the username is not <code>nvidia</code>, create this user on all Sparks:</p>
</blockquote>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-69 > .CodeMirror, .fusion-syntax-highlighter-69 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-69 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_69" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_69" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_69" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sudo useradd -m nvidia
sudo usermod -aG sudo nvidia
sudo passwd nvidia
su &#8211; nvidia</textarea></div><div class="fusion-text fusion-text-80"><blockquote>
<p>These commands: create the user, add them to the sudo group, set a password, and switch to the new user. Ensure the same username exists on all Sparks.</p>
</blockquote>
<blockquote>
<p><strong>If you are using multiple Sparks,</strong> connect the devices to each other with QSFP cables. The cable can be plugged into any of the CX-7 ports on each Spark. To verify the connection, log in to each device with a monitor and keyboard and run the following command:</p>
</blockquote>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-70 > .CodeMirror, .fusion-syntax-highlighter-70 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-70 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_70" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_70" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_70" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">‰·^¿i޵ׯ</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-71 > .CodeMirror, .fusion-syntax-highlighter-71 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-71 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_71" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_71" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_71" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">rocep1s0f0 port 1 ==&gt; enp1s0f0np0 (Down)
rocep1s0f1 port 1 ==&gt; enp1s0f1np1 (Up)
roceP2p1s0f0 port 1 ==&gt; enP2p1s0f0np0 (Down)
roceP2p1s0f1 port 1 ==&gt; enP2p1s0f1np1 (Up)</textarea></div><div class="fusion-text fusion-text-81"><blockquote>
<p>Those shown as <code>(Up)</code> are the ports where the cable is connected. Those shown as <code>(Down)</code> are unused ports. If no interface shows as <code>(Up)</code>, check the QSFP cable and restart the Sparks.</p>
</blockquote>
<p>Ensure your computer is connected to the same network as the Spark. Then, open a terminal on your computer and connect to the Spark via SSH:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-72 > .CodeMirror, .fusion-syntax-highlighter-72 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-72 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_72" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_72" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_72" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">ssh nvidia@</textarea></div><div class="fusion-text fusion-text-82"><p>&nbsp;</p>
<p>On the first connection you will see a fingerprint warning. Type <code>yes</code> and press Enter. Then, when prompted for a password, enter the Spark&#8217;s password:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-73 > .CodeMirror, .fusion-syntax-highlighter-73 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-73 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_73" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_73" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_73" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">The authenticity of host &#8216;192.168.1.148 (192.168.1.148)&#8217; can&#8217;t be established.
ED25519 key fingerprint is SHA256:S6EECYc6Pw2aLoLmhblFZ0QEoeVtJP41jJ5IYsdOmMM.
This key is not known by any other names
Are you sure you want to continue connecting (yes/no/[fingerprint])? yes
Warning: Permanently added &#8216;192.168.1.148&#8217; (ED25519) to the list of known hosts.
nvidia@192.168.1.148&#8217;s password:
Welcome to NVIDIA DGX Spark Version 7.5.0 (GNU/Linux 6.17.0-1026-nvidia aarch64)</p>
<p>System information as of Fri Jul 17 01:26:03 PM UTC 2026</p>
<p>System load: 1.68 Temperature: 50.0 C
Usage of /: 60.6% of 3.67TB Processes: 575
Memory usage: 3% Users logged in: 1
Swap usage: 7% IPv4 address for enP7s7: 192.168.1.148</p>
<p>2 devices have a firmware upgrade available.
Run `fwupdmgr get-upgrades` for more information.</p>
<p>Last login: Fri Jul 17 13:16:30 2026 from 192.168.1.77</textarea></div><div class="fusion-text fusion-text-83"><p>&nbsp;</p>
<p>Once connected, the Spark will begin accepting commands sent from this terminal. Throughout the tutorial, you will enter all commands you encounter into this terminal on your computer.</p>
<hr />
<h3>2. Installing sparkrun</h3>
<p>First, download the uv package manager:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-74 > .CodeMirror, .fusion-syntax-highlighter-74 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-74 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_74" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_74" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_74" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl -LsSf https://astral.sh/uv/install.sh | sh</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-75 > .CodeMirror, .fusion-syntax-highlighter-75 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-75 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_75" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_75" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_75" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">downloading uv 0.11.29 aarch64-unknown-linux-gnu
installing to /home/nvidia/.local/bin
uv
uvx
everything&#8217;s installed!</textarea></div><div class="fusion-text fusion-text-84"><p>&nbsp;</p>
<p>Add uv to your PATH:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-76 > .CodeMirror, .fusion-syntax-highlighter-76 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-76 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_76" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_76" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_76" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">echo &#8216;export PATH=&#8221;$HOME/.local/bin:$PATH&#8221;&#8216; &gt;&gt; ~/.bashrc
source ~/.bashrc</textarea></div><div class="fusion-text fusion-text-85"><p>&nbsp;</p>
<p>PATH updated. The <code>~/.local/bin</code> directory (where uv and sparkrun are installed) is now available and has been added to the PATH for all future terminal sessions. Now install sparkrun:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-77 > .CodeMirror, .fusion-syntax-highlighter-77 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-77 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_77" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_77" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_77" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">uv tool install sparkrun</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-78 > .CodeMirror, .fusion-syntax-highlighter-78 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-78 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_78" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_78" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_78" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Resolved 33 packages in 640ms
Installed 33 packages in 23ms
+ annotated-doc==0.0.4
+ anyio==4.14.2
+ botwinick-utils==0.0.20
+ certifi==2026.6.17
+ click==8.3.3
+ filelock==3.30.2
+ fsspec==2026.6.0
+ h11==0.16.0
+ hf-xet==1.5.2
+ httpcore==1.0.9
+ httpx==0.28.1
+ huggingface-hub==1.8.0
+ idna==3.18
+ linkify-it-py==2.1.0
+ markdown-it-py==4.2.0
+ mdit-py-plugins==0.6.1
+ mdurl==0.1.2
+ packaging==26.2
+ platformdirs==4.10.0
+ pygments==2.20.0
+ python-json-logger==4.1.0
+ pyyaml==6.0.3
+ rich==15.0.0
+ scitrera-app-framework==0.0.69
+ shellingham==1.5.4
+ six==1.17.0
+ sparkrun==0.2.40
+ textual==8.2.5
+ tqdm==4.68.4
+ typer==0.27.0
+ typing-extensions==4.16.0
+ uc-micro-py==2.0.0
+ vpd==0.9.13
Installed 1 executable: sparkrun</textarea></div><div class="fusion-text fusion-text-86"><p>&nbsp;</p>
<p>Let&#8217;s verify the installation:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-79 > .CodeMirror, .fusion-syntax-highlighter-79 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-79 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_79" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_79" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_79" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun &#8211;version</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-80 > .CodeMirror, .fusion-syntax-highlighter-80 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-80 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_80" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_80" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_80" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">sparkrun, version 0.2.40</textarea></div><div class="fusion-text fusion-text-87"><p>&nbsp;</p>
<p>sparkrun v0.2.40 is installed and accessible.</p>
<hr />
<h3>3. Building the Docker Image</h3>
<p>To get the best performance from GPT-OSS 120B, we will build a Docker image that contains CUTLASS MoE and FlashInfer attention kernels compiled specifically for the Spark&#8217;s GPU architecture. This image also includes the <code>--mxfp4-backend</code>, <code>--mxfp4-layers</code> flags and GPT-OSS&#8217;s tiktoken encoding files.</p>
<p>We will build this image using the build-and-copy.sh script from the eugr community project. First, clone the eugr repository:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-81 > .CodeMirror, .fusion-syntax-highlighter-81 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-81 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_81" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_81" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_81" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cd ~
git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker</textarea></div><div class="fusion-text fusion-text-88"><p>&nbsp;</p>
<p>The <code>Dockerfile.mxfp4</code> file in the repository determines the version of the source code that build-and-copy.sh will compile. Before starting the build, we must ensure this version is up to date. Therefore, let&#8217;s query the latest commit hash of the mxfp4_v2 branch in christopherowen&#8217;s community project vLLM fork:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-82 > .CodeMirror, .fusion-syntax-highlighter-82 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-82 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_82" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_82" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_82" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">git ls-remote https://github.com/christopherowen/vllm.git refs/heads/mxfp4_v2 | cut -f1</textarea></div><div class="fusion-text fusion-text-89"><p>&nbsp;</p>
<p>This hash returned by the command will direct build-and-copy.sh to the latest version of the vLLM source code we want. Copy or note it. Then open the Dockerfile.mxfp4 file with nano:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-83 > .CodeMirror, .fusion-syntax-highlighter-83 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-83 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_83" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_83" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_83" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">nano Dockerfile.mxfp4</textarea></div><div class="fusion-text fusion-text-90"><p>&nbsp;</p>
<p>Once the file is open, press Ctrl+W to search within the file, type VLLM_SHA and press Enter. The cursor will land on a line similar to:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-84 > .CodeMirror, .fusion-syntax-highlighter-84 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-84 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_84" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_84" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_84" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">ARG VLLM_SHA=045293d82b832229560ac4a13152a095af603b6e</textarea></div><div class="fusion-text fusion-text-91"><p>&nbsp;</p>
<p>Delete the old hash value and paste the updated hash you noted above. Then press Ctrl+O to save the file, confirm with Enter, and exit nano with Ctrl+X.</p>
<p>Let&#8217;s verify the change we made to the Dockerfile.mxfp4 file:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-85 > .CodeMirror, .fusion-syntax-highlighter-85 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-85 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_85" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_85" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_85" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">grep VLLM_SHA Dockerfile.mxfp4</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-86 > .CodeMirror, .fusion-syntax-highlighter-86 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-86 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_86" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_86" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_86" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">ARG VLLM_SHA=04f641e537e80a67db464c6e65b928dbfab5d647</textarea></div><div class="fusion-text fusion-text-92"><p>&nbsp;</p>
<p>The returned hash value should be the updated hash you copied in the previous step.</p>
<p><strong>Fallback:</strong> If you experience issues with the latest commit, you can use the hash 04f641e537e80a67db464c6e65b928dbfab5d647 (January 28, 2026), which we have tested and confirmed. You can also check for newer commits in the <strong><a style="color: #47d600;" href="https://github.com/christopherowen/vllm/commits/mxfp4_v2">commit history</a> </strong>.</p>
<p>Now build the image:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-87 > .CodeMirror, .fusion-syntax-highlighter-87 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-87 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_87" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_87" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_87" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">DOCKER_BUILDKIT=1 ./build-and-copy.sh &#8211;exp-mxfp4</textarea></div><div class="fusion-text fusion-text-93"><p>&nbsp;</p>
<p>The first build takes approximately 50 minutes; subsequent builds complete in about 5 minutes thanks to caching. When the build completes, you will see the following output:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-88 > .CodeMirror, .fusion-syntax-highlighter-88 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-88 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_88" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_88" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_88" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">=========================================
TIMING STATISTICS
=========================================
Runner Build: 00:50:00
Total Time: 00:50:00
=========================================
Done preparing vllm-node-mxfp4.</textarea></div><div class="fusion-text fusion-text-94"><p>&nbsp;</p>
<p>Verify the image:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-89 > .CodeMirror, .fusion-syntax-highlighter-89 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-89 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_89" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_89" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_89" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker run &#8211;rm &#8211;entrypoint cat vllm-node-mxfp4:latest /workspace/build-metadata.yaml</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-90 > .CodeMirror, .fusion-syntax-highlighter-90 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-90 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_90" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_90" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_90" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/yaml">vllm_commit: 04f641e537e80a67db464c6e65b928dbfab5d647
flashinfer_commit: f349e52496a72a00d8c4ac02c7a1e38523ff7194
gpu_arch: 12.1a
base_image: nvcr.io/nvidia/pytorch:26.01-py3</textarea></div><div class="fusion-text fusion-text-95"><p>&nbsp;</p>
<p>Verify the image&#8217;s vLLM version:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-91 > .CodeMirror, .fusion-syntax-highlighter-91 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-91 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_91" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_91" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_91" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker run &#8211;rm &#8211;entrypoint python3 vllm-node-mxfp4:latest -c &#8220;import vllm; print(vllm.__version__)&#8221;</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-92 > .CodeMirror, .fusion-syntax-highlighter-92 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-92 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_92" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_92" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_92" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">0.1.dev12774+g04f641e53.d20260720</textarea></div><div class="fusion-text fusion-text-96"><p>&nbsp;</p>
<p><strong>If you are using multiple Sparks,</strong> building the image only on the Main Spark is sufficient. sparkrun will automatically synchronize the image to the Workers over the CX-7 network on first launch. No additional step is needed.</p>
<hr />
<h3>4. Setup Wizard</h3>
<p>The sparkrun setup wizard configures all the infrastructure needed to run the Spark with a single interactive command:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-93 > .CodeMirror, .fusion-syntax-highlighter-93 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-93 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_93" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_93" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_93" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun setup wizard</textarea></div><div class="fusion-text fusion-text-97"><p>&nbsp;</p>
<p>Six phases run sequentially, and the phases are interactive. In the first phase, enter <code></code> in the &#8216;Enter host IPs/hostnames&#8217; field and press Enter. For the subsequent prompts, you can simply press Enter (confirming the default options) to continue:</p>
<p><strong>If you are using multiple Sparks,</strong> enter all the Sparks&#8217; IP addresses separated by commas in the &#8216;Enter host IPs/hostnames&#8217; field .</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-94 > .CodeMirror, .fusion-syntax-highlighter-94 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-94 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_94" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_94" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_94" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Welcome to sparkrun 0.2.40 setup wizard!
================================================</p>
<p>Phase 1: Cluster Setup
&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;
Detecting CX7 interfaces on this machine&#8230;
CX7 detected! This machine is a DGX Spark.
Enter host IPs/hostnames (comma-separated) [192.168.1.148]: 192.168.1.148,192.168.1.163
Cluster name [default]: dualspark
SSH username [nvidia]:
Created cluster &#8216;dualspark&#8217; with 2 host(s), set as default.</textarea></div><div class="fusion-text fusion-text-98"><p>&nbsp;</p>
<p>The second phase sets up passwordless SSH. Type <code>Y</code> and press Enter:</p>
<p><strong>If you are using multiple Sparks,</strong> the wizard connects to the Worker Sparks via SSH during this phase and asks for the password. Enter the password — this step is done once; after the wizard distributes the passwordless SSH keys, no password will be asked for in subsequent connections.</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-95 > .CodeMirror, .fusion-syntax-highlighter-95 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-95 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_95" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_95" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_95" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Phase 2: SSH Mesh
&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;
Set up SSH mesh across 2 host(s) + this machine? [Y/n]: Y
=== Phase 1: Connectivity check ===[*] Checking SSH connectivity to nvidia@192.168.1.148 &#8230;[*] Checking SSH connectivity to nvidia@192.168.1.163 &#8230;</p>
<p>=== Phase 4: Install keys so every host trusts every other host ===[*] Installing key from 192.168.1.148 onto all other hosts &#8230;
&#8211; 192.168.1.148 -&gt; 192.168.1.163[*] Installing key from 192.168.1.163 onto all other hosts &#8230;
&#8211; 192.168.1.163 -&gt; 192.168.1.148</p>
<p>=== Done ===
All hosts should now be able to SSH to each other as &#8216;nvidia&#8217; without passwords.</p>
<p>Detecting management IPs on cluster hosts&#8230;
Updating cluster hosts to management IPs:
192.168.1.148 -&gt; 127.0.0.1 (local)
Cluster &#8216;dualspark&#8217; updated.</textarea></div><div class="fusion-text fusion-text-99"><p>&nbsp;</p>
<p>The third phase configures the CX-7 high-speed network interfaces. It is automatically skipped if a single Spark is used. Type <code>Y</code> and press Enter:</p>
<p><strong>If you are using multiple Sparks,</strong> the wizard automatically detects CX-7 interfaces, selects non-overlapping subnets, assigns a static IP to each Spark, sets MTU 9000 (jumbo frame), and writes the <code>/etc/netplan/40-cx7.yaml</code> file and applies it with <code>netplan apply</code>. This operation requires root privileges — enter the sudo password.</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-96 > .CodeMirror, .fusion-syntax-highlighter-96 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-96 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_96" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_96" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_96" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Phase 3: CX7 Network Configuration
&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;
Configures high-speed CX7 networking between hosts.
Configure CX7 networking? [Y/n]: Y
Topology: switch
Subnets: 192.168.0.0/24, 192.168.2.0/24[sudo] password for nvidia: </textarea></div><div class="fusion-text fusion-text-100"><p>&nbsp;</p>
<p>After CX-7 configuration, the wizard automatically refreshes the SSH network with the new IPs. This phase requires no additional input.</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-97 > .CodeMirror, .fusion-syntax-highlighter-97 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-97 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_97" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_97" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_97" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Phase 3b: Re-meshing SSH after CX7 IP changes
&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;
CX7 configuration changed network IPs. Re-running SSH mesh
to ensure full connectivity across all interfaces.</p>
<p>Verifying SSH connectivity to cluster hosts&#8230;
All 2 host(s) reachable.</p>
<p>Checking reachability from control machine&#8230;
Reachable: 127.0.0.1, 192.168.1.163, 192.168.0.148, 192.168.2.148, 192.168.0.163, 192.168.2.163</textarea></div><div class="fusion-text fusion-text-101"><p>&nbsp;</p>
<p>The fourth phase verifies that the user is in the docker group. Type <code>Y</code> and press Enter:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-98 > .CodeMirror, .fusion-syntax-highlighter-98 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-98 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_98" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_98" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_98" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Phase 4: Docker Group Membership
&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;
Ensures user can run Docker commands without sudo.
Add &#8216;nvidia&#8217; to the docker group on all hosts? [Y/n]: Y
127.0.0.1: &#8216;nvidia&#8217; already a member
192.168.1.163: &#8216;nvidia&#8217; already a member</textarea></div><div class="fusion-text fusion-text-102"><p>&nbsp;</p>
<p>The fifth phase installs restricted sudoers rules. The installed rules are for passwordlessly fixing HuggingFace cache ownership and passwordlessly clearing the Linux page cache. Broad sudo privileges are not granted. Type <code>Y</code> and press Enter:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-99 > .CodeMirror, .fusion-syntax-highlighter-99 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-99 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_99" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_99" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_99" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Phase 5: Sudoers Entries
&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;
Scoped sudoers for fix-permissions + clear-cache (no broad sudo).
Install sudoers entries? [Y/n]: Y
fix-permissions: 2/2 host(s)
clear-cache: 2/2 host(s)</textarea></div><div class="fusion-text fusion-text-103"><p>&nbsp;</p>
<p>The sixth phase installs earlyoom OOM protection. earlyoom terminates inference processes instead of locking up the system when memory limits are approached in the DGX Spark&#8217;s unified memory architecture. Type<strong> <code>Y</code></strong> and press Enter:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-100 > .CodeMirror, .fusion-syntax-highlighter-100 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-100 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_100" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_100" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_100" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Phase 6: earlyoom OOM Protection
&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;
Prevents system hangs by proactively managing memory pressure.
Install earlyoom? [Y/n]: Y
earlyoom configured on 2/2 host(s).</p>
<p>Setup Complete!
================================================</p>
<p>Cluster: dualspark (2 hosts, set as default)
SSH mesh: OK
CX7: configured (switch)
SSH remesh: OK
Docker: OK (2/2)
Sudoers: installed (fix-permissions, clear-cache)
earlyoom: installed</textarea></div><div class="fusion-text fusion-text-104"><hr />
<h3>5. Creating the Recipe</h3>
<p>A recipe is a YAML file that defines how sparkrun runs the model. The model, Docker image, vLLM flags, and memory settings are consolidated in a single file. This recipe can be used for both a single Spark (TP=1) and multiple Sparks (TP&gt;1) — the only difference is that the <code>--tp</code> value in the launch command is adjusted according to the number of Sparks.</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-101 > .CodeMirror, .fusion-syntax-highlighter-101 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-101 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_101" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_101" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_101" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cat &gt; ~/gpt-oss-120b-eugr-mxfp4.yaml &lt;&lt; &#8216;RECIPE&#8217;
# GPT-OSS 120B MXFP4 — eugr &#8211;exp-mxfp4 build, CUTLASS MoE, FLASHINFER attention
# Usage:
# sparkrun run ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;rootful # single node (TP=1)
# sparkrun run ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;rootful &#8211;tp 2 # dual node (TP=2)
recipe_version: &#8220;2&#8221;
model: openai/gpt-oss-120b
runtime: vllm
container: vllm-node-mxfp4</p>
<p>metadata:
description: GPT-OSS 120B MXFP4 — CUTLASS MoE, FLASHINFER attention, eugr &#8211;exp-mxfp4 build</p>
<p>defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.7
max_num_batched_tokens: 8192</p>
<p>env:
VLLM_USE_FLASHINFER_MOE_MXFP4_MXFP8: &#8220;1&#8221;</p>
<p>command: |
vllm serve {model}
&#8211;tool-call-parser openai
&#8211;reasoning-parser openai_gptoss
&#8211;enable-auto-tool-choice
&#8211;tensor-parallel-size {tensor_parallel}
&#8211;distributed-executor-backend ray
&#8211;gpu-memory-utilization {gpu_memory_utilization}
&#8211;enable-prefix-caching
&#8211;load-format fastsafetensors
&#8211;quantization mxfp4
&#8211;mxfp4-backend CUTLASS
&#8211;mxfp4-layers moe,qkv,o,lm_head
&#8211;attention-backend FLASHINFER
&#8211;kv-cache-dtype fp8
&#8211;max-num-batched-tokens {max_num_batched_tokens}
&#8211;host {host}
&#8211;port {port}
RECIPE</textarea></div><div class="fusion-text fusion-text-105"><p>&nbsp;</p>
<p>Verify the file:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-102 > .CodeMirror, .fusion-syntax-highlighter-102 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-102 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_102" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_102" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_102" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cat ~/gpt-oss-120b-eugr-mxfp4.yaml</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-103 > .CodeMirror, .fusion-syntax-highlighter-103 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-103 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_103" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_103" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_103" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/yaml"># GPT-OSS 120B MXFP4 — eugr &#8211;exp-mxfp4 build, CUTLASS MoE, FLASHINFER attention
# Usage:
# sparkrun run ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;rootful # single node (TP=1)
# sparkrun run ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;rootful &#8211;tp 2 # dual node (TP=2)
recipe_version: &#8220;2&#8221;
model: openai/gpt-oss-120b
runtime: vllm
container: vllm-node-mxfp4</p>
<p>metadata:
description: GPT-OSS 120B MXFP4 — CUTLASS MoE, FLASHINFER attention, eugr &#8211;exp-mxfp4 build</p>
<p>defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.7
max_num_batched_tokens: 8192</p>
<p>env:
VLLM_USE_FLASHINFER_MOE_MXFP4_MXFP8: &#8220;1&#8221;</p>
<p>command: |
vllm serve {model}
&#8211;tool-call-parser openai
&#8211;reasoning-parser openai_gptoss
&#8211;enable-auto-tool-choice
&#8211;tensor-parallel-size {tensor_parallel}
&#8211;distributed-executor-backend ray
&#8211;gpu-memory-utilization {gpu_memory_utilization}
&#8211;enable-prefix-caching
&#8211;load-format fastsafetensors
&#8211;quantization mxfp4
&#8211;mxfp4-backend CUTLASS
&#8211;mxfp4-layers moe,qkv,o,lm_head
&#8211;attention-backend FLASHINFER
&#8211;kv-cache-dtype fp8
&#8211;max-num-batched-tokens {max_num_batched_tokens}
&#8211;host {host}
&#8211;port {port}</textarea></div><div class="fusion-text fusion-text-106"><p>&nbsp;</p>
<p>Key fields in the recipe and the reasons they were selected:</p>
<ul>
<li><strong><code>container</code></strong>
<ul>
<li><strong>Value:</strong> <code>vllm-node-mxfp4</code></li>
<li><strong>Why:</strong> Image built in step 3 — includes the CUTLASS/FlashInfer forks</li>
</ul>
</li>
<li><strong><code>tensor_parallel</code></strong>
<ul>
<li><strong>Value:</strong> <code>1</code> (default)</li>
<li><strong>Why:</strong> Default for a single Spark; overridden with the <code>--tp</code> value when using multiple Sparks</li>
</ul>
</li>
<li><strong><code>gpu_memory_utilization</code></strong>
<ul>
<li><strong>Value:</strong> <code>0.7</code></li>
<li><strong>Why:</strong> Safe value for DGX Spark UMA</li>
</ul>
</li>
<li><strong><code>--distributed-executor-backend ray</code></strong>
<ul>
<li><strong>Value:</strong> Ray</li>
<li><strong>Why:</strong> sparkrun automatically detects Ray and sets up the Ray cluster. For TP=1, Ray runs on a single node.</li>
</ul>
</li>
<li><strong><code>--mxfp4-backend CUTLASS</code></strong>
<ul>
<li><strong>Value:</strong> CUTLASS</li>
<li><strong>Why:</strong> Custom CUTLASS MXFP4 MoE GEMM kernel</li>
</ul>
</li>
<li><strong><code>--mxfp4-layers moe,qkv,o,lm_head</code></strong>
<ul>
<li><strong>Value:</strong> Full</li>
<li><strong>Why:</strong> All layers are quantized to FP4</li>
</ul>
</li>
<li><strong><code>--attention-backend FLASHINFER</code></strong>
<ul>
<li><strong>Value:</strong> FlashInfer</li>
<li><strong>Why:</strong> Custom FlashInfer kernel supporting GPT-OSS attention architecture</li>
</ul>
</li>
<li><strong><code>--kv-cache-dtype fp8</code></strong>
<ul>
<li><strong>Value:</strong> FP8</li>
<li><strong>Why:</strong> Reduced memory usage for the KV cache</li>
</ul>
</li>
<li><strong><code>--load-format fastsafetensors</code></strong>
<ul>
<li><strong>Value:</strong> Fast loader</li>
<li><strong>Why:</strong> ~41-second model loading time</li>
</ul>
</li>
<li><strong><code>--reasoning-parser openai_gptoss</code></strong>
<ul>
<li><strong>Value:</strong> Harmony</li>
<li><strong>Why:</strong> Separates the reasoning process into the <code>reasoning</code> field</li>
</ul>
</li>
<li><strong><code>--tool-call-parser openai</code></strong>
<ul>
<li><strong>Value:</strong> Tool usage</li>
<li><strong>Why:</strong> Required for <code>--enable-auto-tool-choice</code></li>
</ul>
</li>
</ul>
<p><strong>Important:</strong><br />The <code>--rootful</code> flag is not specified in the recipe; it is passed in the launch command.<br />FlashInfer compiles CUTLASS attention kernels at runtime, and this compilation requires root privileges.<br />The launch commands are shown in the “Launch the Model” step of the relevant section.</p>
</div><div class="fusion-text fusion-text-107"><h3>6. Pre-Launch Checks</h3>
<p>Verify that `sparkrun` correctly parses the recipe and that the memory budget is appropriate:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-104 > .CodeMirror, .fusion-syntax-highlighter-104 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-104 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_104" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_104" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_104" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun show ~/gpt-oss-120b-eugr-mxfp4.yaml</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-105 > .CodeMirror, .fusion-syntax-highlighter-105 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-105 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_105" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_105" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_105" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Name: /home/nvidia/gpt-oss-120b-eugr-mxfp4.yaml
Description: GPT-OSS 120B MXFP4 — CUTLASS MoE, FLASHINFER attention, eugr &#8211;exp-mxfp4 build
Runtime: vllm-ray
Model: openai/gpt-oss-120b
Container: vllm-node-mxfp4
Nodes: 1 &#8211; unlimited</p>
<p>Defaults:
gpu_memory_utilization: 0.7
host: 0.0.0.0
max_num_batched_tokens: 8192
port: 8000
tensor_parallel: 1</p>
<p>Environment:
VLLM_USE_FLASHINFER_MOE_MXFP4_MXFP8=1</p>
<p>Command:
vllm serve {model}
&#8211;tool-call-parser openai
&#8211;reasoning-parser openai_gptoss
&#8211;enable-auto-tool-choice
&#8211;tensor-parallel-size {tensor_parallel}
&#8211;distributed-executor-backend ray
&#8211;gpu-memory-utilization {gpu_memory_utilization}
&#8211;enable-prefix-caching
&#8211;load-format fastsafetensors
&#8211;quantization mxfp4
&#8211;mxfp4-backend CUTLASS
&#8211;mxfp4-layers moe,qkv,o,lm_head
&#8211;attention-backend FLASHINFER
&#8211;kv-cache-dtype fp8
&#8211;max-num-batched-tokens {max_num_batched_tokens}
&#8211;host {host}
&#8211;port {port}
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.</p>
<p>VRAM Estimation:
Model dtype: mxfp4
KV cache dtype: bfloat16
Architecture: 36 layers, 8 KV heads, 64 head_dim
Model weights: 60.77 GB
Tensor parallel: 1
Per-GPU total: 60.77 GB
DGX Spark fit: YES</p>
<p>GPU Memory Budget:
gpu_memory_utilization: 70%
Usable GPU memory: 84.7 GB (121 GB x 70%)
Available for KV: 23.9 GB
Max context tokens: 348,538</textarea></div><div class="fusion-text fusion-text-108"><p>&nbsp;</p>
<p>sparkrun correctly parsed the recipe, selected vllm-ray, and gave the &#8216;DGX Spark fit: YES&#8217; confirmation.</p>
<p>Now preview the launch plan:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-106 > .CodeMirror, .fusion-syntax-highlighter-106 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-106 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_106" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_106" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_106" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun run ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;rootful &#8211;dry-run</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-107 > .CodeMirror, .fusion-syntax-highlighter-107 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-107 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_107" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_107" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_107" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">sparkrun v0.2.40</p>
<p>Runtime: vllm-ray
Image: vllm-node-mxfp4
Model: openai/gpt-oss-120b
Mode: solo
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.</p>
<p>VRAM Estimation:
Model dtype: mxfp4
KV cache dtype: bfloat16
Architecture: 36 layers, 8 KV heads, 64 head_dim
Model weights: 60.77 GB
Tensor parallel: 1
Per-GPU total: 60.77 GB
DGX Spark fit: YES</p>
<p>GPU Memory Budget:
gpu_memory_utilization: 70%
Usable GPU memory: 84.7 GB (121 GB x 70%)
Available for KV: 23.9 GB
Max context tokens: 348,538</p>
<p>Hosts: default cluster &#8216;dualspark&#8217;
Target: 127.0.0.1</p>
[1/6] Preparing
done (0.0s)[2/6] Building image — skipped (no builder)[3/6] Distributing resources
done (0.1s)[4/6] Syncing tuning configs
done (0.0s)[5/6] Launching vllm runtime
Step 1/3: Detecting InfiniBand
Step 2/3: Launching container
Step 3/3: Executing serve command
done (0.0s)
Cluster: sparkrun_531d9a939d23</p>
<p>Serve command:
vllm serve openai/gpt-oss-120b
&#8211;tool-call-parser openai
&#8211;reasoning-parser openai_gptoss
&#8211;enable-auto-tool-choice
&#8211;tensor-parallel-size 1
&#8211;distributed-executor-backend ray
&#8211;gpu-memory-utilization 0.7
&#8211;enable-prefix-caching
&#8211;load-format fastsafetensors
&#8211;quantization mxfp4
&#8211;mxfp4-backend CUTLASS
&#8211;mxfp4-layers moe,qkv,o,lm_head
&#8211;attention-backend FLASHINFER
&#8211;kv-cache-dtype fp8
&#8211;max-num-batched-tokens 8192
&#8211;host 0.0.0.0
&#8211;port 8000</p>
[6/6] Post-launch hooks — skipped</textarea></div><div class="fusion-text fusion-text-109"><p>&nbsp;</p>
<p>Dry run successful. The recipe is valid, the serve command includes the flags we added, and it is ready to launch. If the model is not installed on the Spark, sparkrun will automatically download it from Hugging Face on first launch. The GPT-OSS 120B model is approximately 183 GB; download time depends on your internet speed. The dry run does not trigger this download — downloading only happens on the actual launch.</p>
<p><strong>If you are using multiple Sparks,</strong> set the <code>--tp</code> value according to the number of Sparks for the TP dry-run (for example, <code>--tp 2</code> for two Sparks). You will see <code>--tensor-parallel-size 2</code> in the serve command, and the mode will appear as <code>cluster (2 nodes)</code> instead of <code>solo</code>:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-108 > .CodeMirror, .fusion-syntax-highlighter-108 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-108 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_108" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_108" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_108" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun run ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;rootful &#8211;tp 2 &#8211;dry-run</textarea></div><div class="fusion-text fusion-text-110"><hr />
<h2>Running with a Single Spark</h2>
<h3>1. Clearing the Cache</h3>
<p>Before starting vLLM, clear the filesystem cache:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-109 > .CodeMirror, .fusion-syntax-highlighter-109 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-109 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_109" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_109" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_109" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sudo sh -c &#8216;sync; echo 3 &gt; /proc/sys/vm/drop_caches&#8217;</textarea></div><div class="fusion-text fusion-text-111"><p>&nbsp;</p>
<p>The main reason for doing this is the DGX Spark&#8217;s unified memory architecture: The operating system caches the model files it reads from disk in RAM. vLLM loads the model weights from here into GPU memory. After loading is complete, the data remaining in the cache is not used again but is not cleaned up immediately. In systems with separate memory, this is not significant. Since inference runs in GPU memory, RAM fullness does not affect performance. On the Spark, however, since the CPU and GPU share the same RAM, the cache reduces the space available to the GPU. The command clears this cache, providing maximum memory for vLLM.</p>
<p>When prompted for a password, enter the Spark&#8217;s password. The command produces no output; it completes silently:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-110 > .CodeMirror, .fusion-syntax-highlighter-110 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-110 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_110" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_110" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_110" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">[sudo] password for nvidia:</textarea></div><div class="fusion-text fusion-text-112"><hr />
<h3>2. Launching the Model</h3>
<p>Now let&#8217;s launch the model:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-111 > .CodeMirror, .fusion-syntax-highlighter-111 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-111 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_111" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_111" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_111" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun run ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;rootful &#8211;no-follow</textarea></div><div class="fusion-text fusion-text-113"><p>&nbsp;</p>
<p>The <code>--rootful</code> flag runs the container with root privileges. FlashInfer compiles CUTLASS attention kernels at runtime and requires root privileges for this compilation. The <code>--no-follow</code> flag causes sparkrun to return to the command line after launching the containers.</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-112 > .CodeMirror, .fusion-syntax-highlighter-112 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-112 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_112" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_112" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_112" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Note: 1 nodes required, using 1 of 2 hosts
sparkrun v0.2.40</p>
<p>Runtime: vllm-ray
Image: vllm-node-mxfp4
Model: openai/gpt-oss-120b
Mode: solo</p>
<p>VRAM Estimation:
Model dtype: mxfp4
KV cache dtype: bfloat16
Architecture: 36 layers, 8 KV heads, 64 head_dim
Model weights: 60.77 GB
Tensor parallel: 1
Per-GPU total: 60.77 GB
DGX Spark fit: YES</p>
<p>GPU Memory Budget:
gpu_memory_utilization: 70%
Usable GPU memory: 84.7 GB (121 GB x 70%)
Available for KV: 23.9 GB
Max context tokens: 348,538</p>
<p>Hosts: default cluster &#8216;dualspark&#8217;
Target: 127.0.0.1</p>
[1/6] Preparing
done (0.0s)[2/6] Building image — skipped (no builder)[3/6] Distributing resources
Fetching 37 files: 100%|██████████| 37/37 [00:00&lt;00:00, 10007.69it/s]
done (0.6s)[4/6] Syncing tuning configs
done (0.0s)[5/6] Launching vllm runtime
Step 1/3: Detecting InfiniBand
Step 2/3: Launching container
Step 3/3: Executing serve command
done (7.0s)
Cluster: sparkrun_531d9a939d23</p>
<p>Serve command:
vllm serve openai/gpt-oss-120b
&#8211;tool-call-parser openai
&#8211;reasoning-parser openai_gptoss
&#8211;enable-auto-tool-choice
&#8211;tensor-parallel-size 1
&#8211;distributed-executor-backend ray
&#8211;gpu-memory-utilization 0.7
&#8211;enable-prefix-caching
&#8211;load-format fastsafetensors
&#8211;quantization mxfp4
&#8211;mxfp4-backend CUTLASS
&#8211;mxfp4-layers moe,qkv,o,lm_head
&#8211;attention-backend FLASHINFER
&#8211;kv-cache-dtype fp8
&#8211;max-num-batched-tokens 8192
&#8211;host 0.0.0.0
&#8211;port 8000</p>
<p>Runtime versions:
cuda: 13.1
nccl: (2, 29, 2)
python: 3.12.3
torch: 2.10.0a0+a36e1d39eb.nv26.01.42222806
vllm: 0.1.dev12774+g04f641e53.d20260720</p>
[6/6] Post-launch hooks — skipped</textarea></div><div class="fusion-text fusion-text-114"><p>&nbsp;</p>
<p>sparkrun successfully completed all 6 steps. Image building was skipped, and the image we built during setup was used. Since the model files were already on disk, distribution completed quickly. All flags were correctly resolved within the serve command.</p>
<hr />
<h3>3. Monitoring the Startup Process</h3>
<p>After the model download completes, it may take a few minutes for vLLM to become ready to serve. During this process, vLLM loads the model weights into GPU memory, compiles GPU kernels, and allocates memory for inference.</p>
<p>To monitor the vLLM logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-113 > .CodeMirror, .fusion-syntax-highlighter-113 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-113 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_113" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_113" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_113" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun logs ~/gpt-oss-120b-eugr-mxfp4.yaml</textarea></div><div class="fusion-text fusion-text-115"><p>&nbsp;</p>
<p>You will see the following in the logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-114 > .CodeMirror, .fusion-syntax-highlighter-114 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-114 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_114" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_114" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_114" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">(APIServer pid=107) vLLM API server version 0.1.dev12774+g04f641e53.d20260720
(APIServer pid=107) non-default args: {&#8216;quantization&#8217;: &#8216;mxfp4&#8217;, &#8216;mxfp4_backend&#8217;: &#8216;CUTLASS&#8217;,
&#8216;mxfp4_layers&#8217;: &#8216;moe,qkv,o,lm_head&#8217;, &#8216;attention_backend&#8217;: &#8216;FLASHINFER&#8217;,
&#8216;distributed_executor_backend&#8217;: &#8216;ray&#8217;, &#8216;tensor_parallel_size&#8217;: 1, &#8230;}</p>
<p>(EngineCore_DP0 pid=512) Started a local Ray instance.</p>
<p>(RayWorkerWrapper pid=1347) SM12x detected &#8211; using native FlashInfer CUTLASS attention
instead of TRT-LLM attention (cubins not available for SM12x)
(RayWorkerWrapper pid=1347) Using AttentionBackendEnum.FLASHINFER backend.
(RayWorkerWrapper pid=1347) [MXFP4] Using backend: CUTLASS (&#8211;mxfp4-backend)
(RayWorkerWrapper pid=1347) Loading safetensors: 100%|██████████| 15/15 [00:41&lt;00:00] (RayWorkerWrapper pid=1347) Loading weights took 41.43 seconds (RayWorkerWrapper pid=1347) [MXFP4] lm_head quantized: torch.Size([201088, 2880]) BF16 -&gt; torch.Size([201088, 1440]) FP4 (4x smaller)
(RayWorkerWrapper pid=1347) Model loading took 61.33 GiB memory</p>
<p>(RayWorkerWrapper pid=1347) torch.compile takes 112.10 s in total
(RayWorkerWrapper pid=1347) Available KV cache memory: 8.460000 GiB
(EngineCore_DP0 pid=512) GPU KV cache size: 246,416 tokens
(EngineCore_DP0 pid=512) Maximum concurrency for 131,072 tokens per request: 3.54x</p>
<p>(APIServer pid=107) Application startup complete.</textarea></div><div class="fusion-text fusion-text-116"><p>&nbsp;</p>
<p>If you see the <code>Application startup complete.</code> line, the model server is ready. Press <code>Ctrl+C</code> to stop watching the logs. The server will continue running in the background.</p>
<hr />
<h3>4. Testing the Model</h3>
<p>First, verify that the server is running:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-115 > .CodeMirror, .fusion-syntax-highlighter-115 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-115 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_115" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_115" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_115" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl -s -o /dev/null -w &#8220;HTTP %{http_code}&#8221; http://:8000/health</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-116 > .CodeMirror, .fusion-syntax-highlighter-116 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-116 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_116" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_116" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_116" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">HTTP 200</textarea></div><div class="fusion-text fusion-text-117"><p>&nbsp;</p>
<p>An &#8216;HTTP 200&#8217; response indicates the server is healthy. Now check the sparkrun container status:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-117 > .CodeMirror, .fusion-syntax-highlighter-117 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-117 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_117" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_117" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_117" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun status</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-118 > .CodeMirror, .fusion-syntax-highlighter-118 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-118 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_118" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_118" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_118" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Job: /home/nvidia/gpt-oss-120b-eugr-mxfp4.yaml (tp=1) [531d9a939d23] (1 container(s))
node_0 127.0.0.1 Up 2 minutes vllm-node-mxfp4</p>
<p>Total: 1 container(s) across 1 host(s)</textarea></div><div class="fusion-text fusion-text-118"><p>&nbsp;</p>
<p>List the registered models:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-119 > .CodeMirror, .fusion-syntax-highlighter-119 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-119 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_119" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_119" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_119" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl -s http://localhost:8000/v1/models | python3 -m json.tool</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-120 > .CodeMirror, .fusion-syntax-highlighter-120 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-120 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_120" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_120" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_120" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="application/json">{
&#8220;object&#8221;: &#8220;list&#8221;,
&#8220;data&#8221;: [
{
&#8220;id&#8221;: &#8220;openai/gpt-oss-120b&#8221;,
&#8220;object&#8221;: &#8220;model&#8221;,
&#8220;created&#8221;: 1784286672,
&#8220;owned_by&#8221;: &#8220;vllm&#8221;,
&#8220;root&#8221;: &#8220;openai/gpt-oss-120b&#8221;,
&#8220;parent&#8221;: null,
&#8220;max_model_len&#8221;: 131072
}
]
}</textarea></div><div class="fusion-text fusion-text-119"><p>&nbsp;</p>
<p>The model <code>openai/gpt-oss-120b</code> is registered.</p>
<p>Now test the model with the prompt &#8220;What is 12*17?&#8221;:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-121 > .CodeMirror, .fusion-syntax-highlighter-121 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-121 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_121" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_121" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_121" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl -s http://localhost:8000/v1/chat/completions
-H &#8220;Content-Type: application/json&#8221;
-d &#8216;{
&#8220;model&#8221;: &#8220;openai/gpt-oss-120b&#8221;,
&#8220;messages&#8221;: [{&#8220;role&#8221;: &#8220;user&#8221;, &#8220;content&#8221;: &#8220;What is 12*17?&#8221;}],
&#8220;max_tokens&#8221;: 200
}&#8217; | python3 -m json.tool</textarea></div><div class="fusion-text fusion-text-120"><p>&nbsp;</p>
<p>A simplified version of the response is below. In the <code>content</code> field (the model&#8217;s answer) you will see <code>204</code>. From this, you can tell the model performed the multiplication correctly. Additionally, the response has a <code>reasoning</code> field. This field contains the model&#8217;s reasoning process. GPT-OSS&#8217;s Harmony format is working: the reasoning process and the final answer are presented in separate fields:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-122 > .CodeMirror, .fusion-syntax-highlighter-122 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-122 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_122" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_122" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_122" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="application/json">{
&#8220;id&#8221;: &#8220;chatcmpl-b863d6c512722c38&#8221;,
&#8220;model&#8221;: &#8220;openai/gpt-oss-120b&#8221;,
&#8220;choices&#8221;: [
{
&#8220;index&#8221;: 0,
&#8220;message&#8221;: {
&#8220;role&#8221;: &#8220;assistant&#8221;,
&#8220;content&#8221;: &#8220;(12 times 17 = 204)&#8221;,
&#8220;reasoning&#8221;: &#8220;User asks a simple multiplication: 12*17 = 204. Provide answer.&#8221;,
&#8220;reasoning_content&#8221;: &#8220;User asks a simple multiplication: 12*17 = 204. Provide answer.&#8221;
},
&#8220;finish_reason&#8221;: &#8220;stop&#8221;
}
],
&#8220;usage&#8221;: {
&#8220;prompt_tokens&#8221;: 72,
&#8220;total_tokens&#8221;: 110,
&#8220;completion_tokens&#8221;: 38
}
}</textarea></div><div class="fusion-text fusion-text-121"><p>&nbsp;</p>
<p>vLLM is now running and serving on the Spark&#8217;s port 8000. We can ask GPT-OSS 120B questions and get answers.</p>
<hr />
<h3>5. Shutdown</h3>
<p>When you are done, stop the model:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-123 > .CodeMirror, .fusion-syntax-highlighter-123 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-123 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_123" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_123" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_123" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun stop ~/gpt-oss-120b-eugr-mxfp4.yaml</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-124 > .CodeMirror, .fusion-syntax-highlighter-124 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-124 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_124" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_124" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_124" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Workload stopped on 1 host(s).</textarea></div><div class="fusion-text fusion-text-122"><p>&nbsp;</p>
<p>This command stops the container and frees up memory. However, the container, the Docker image, and the model files remain on disk. This means you don&#8217;t need to re-download to start again. Simply re-run the <code>sparkrun run</code> command from step 2.</p>
<hr />
<h2>Running with Two Sparks</h2>
<h3>1. Clearing the Cache</h3>
<p>Before starting vLLM, clear the filesystem cache on both Sparks. The main reason for doing this is the DGX Spark&#8217;s unified memory architecture (UMA): The operating system caches the model files it reads from disk in RAM. vLLM loads the model weights from here into GPU memory. After loading is complete, the data remaining in the cache is not used again but is not cleaned up immediately. In systems with separate memory, this is not significant. Since inference runs in GPU memory, RAM fullness does not affect performance. On the Spark, however, since the CPU and GPU share the same RAM, the cache reduces the space available to the GPU. The command clears this cache, providing maximum memory for vLLM.</p>
<p>Clear the cache on the Main Spark:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-125 > .CodeMirror, .fusion-syntax-highlighter-125 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-125 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_125" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_125" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_125" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sudo sh -c &#8216;sync; echo 3 &gt; /proc/sys/vm/drop_caches&#8217;</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-126 > .CodeMirror, .fusion-syntax-highlighter-126 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-126 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_126" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_126" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_126" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">[sudo] password for nvidia:</textarea></div><div class="fusion-text fusion-text-123"><p>&nbsp;</p>
<p>Apply the same cleanup on the Worker Spark via SSH from the Main Spark:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-127 > .CodeMirror, .fusion-syntax-highlighter-127 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-127 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_127" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_127" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_127" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">ssh nvidia@ &#8220;sudo sh -c &#8216;sync; echo 3 &gt; /proc/sys/vm/drop_caches'&#8221;</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-128 > .CodeMirror, .fusion-syntax-highlighter-128 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-128 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_128" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_128" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_128" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">[sudo] password for nvidia:</textarea></div><div class="fusion-text fusion-text-124"><hr />
<h3>2. Launching the Model</h3>
<p>Now, let&#8217;s launch the model on two Sparks:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-129 > .CodeMirror, .fusion-syntax-highlighter-129 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-129 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_129" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_129" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_129" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun run ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;rootful &#8211;tp 2 &#8211;no-follow</textarea></div><div class="fusion-text fusion-text-125"><p>&nbsp;</p>
<p>The <code>--tp 2</code> flag tells sparkrun to run the model split (tensor parallel) across two Sparks. The <code>--rootful</code> flag runs the container with root privileges — FlashInfer compiles CUTLASS attention kernels at runtime and requires root privileges for this compilation. The <code>--no-follow</code> flag causes sparkrun to return to the command line after launching the containers. sparkrun automatically synchronizes the image to the Worker (skips if the same ID), copies the model to the Worker (skips if already present), sets up the Ray cluster, configures NCCL for the CX-7 interfaces, and launches containers on both Sparks.</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-130 > .CodeMirror, .fusion-syntax-highlighter-130 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-130 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_130" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_130" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_130" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/yaml">sparkrun v0.2.40&lt;/p&gt;
&lt;p&gt;Runtime: vllm-ray&lt;br /&gt;
Image: vllm-node-mxfp4&lt;br /&gt;
Model: openai/gpt-oss-120b&lt;br /&gt;
Mode: cluster (2 nodes)&lt;br /&gt;
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.&lt;br /&gt;
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.&lt;/p&gt;
&lt;p&gt;VRAM Estimation:&lt;br /&gt;
Model dtype: mxfp4&lt;br /&gt;
KV cache dtype: bfloat16&lt;br /&gt;
Architecture: 36 layers, 8 KV heads, 64 head_dim&lt;br /&gt;
Model weights: 60.77 GB&lt;br /&gt;
Tensor parallel: 2&lt;br /&gt;
Per-GPU total: 30.38 GB&lt;br /&gt;
DGX Spark fit: YES&lt;/p&gt;
&lt;p&gt;GPU Memory Budget:&lt;br /&gt;
gpu_memory_utilization: 70%&lt;br /&gt;
Usable GPU memory: 84.7 GB (121 GB x 70%)&lt;br /&gt;
Available for KV: 54.3 GB&lt;br /&gt;
Max context tokens: 1,582,071&lt;/p&gt;
&lt;p&gt;Hosts: default cluster &#8216;dualspark'&lt;br /&gt;
Head: 127.0.0.1&lt;br /&gt;
Workers: 192.168.1.163&lt;/p&gt;[1/6] Preparing&lt;br /&gt;
done (0.0s)[2/6] Building image — skipped (no builder)[3/6] Distributing resources&lt;br /&gt;
Distributing image vllm-node-mxfp4 to 2 host(s)&lt;br /&gt;
Container image stale on 1 of 2 host(s), syncing&lt;br /&gt;
Distributing model openai/gpt-oss-120b to 2 host(s)&lt;br /&gt;
Fetching 37 files: 100%|██████████| 37/37 [00:00&amp;lt;00:00, 7261.00it/s]
Model synced to 2 host(s)&lt;br /&gt;
done (187.3s)[4/6] Syncing tuning configs&lt;br /&gt;
done (0.0s)[5/6] Launching vllm runtime&lt;br /&gt;
Step 1/5: Cleaning up existing containers&lt;br /&gt;
Step 2/5: Detecting InfiniBand&lt;br /&gt;
Step 3/5: Launching Ray head&lt;br /&gt;
Step 4/5: Launching Ray workers&lt;br /&gt;
Step 5/5: Executing serve command on head&lt;br /&gt;
done (17.9s)&lt;br /&gt;
Cluster: sparkrun_631038688bb2&lt;/p&gt;
&lt;p&gt;Serve command:&lt;br /&gt;
vllm serve openai/gpt-oss-120b &lt;br /&gt;
&#8211;tool-call-parser openai &lt;br /&gt;
&#8211;reasoning-parser openai_gptoss &lt;br /&gt;
&#8211;enable-auto-tool-choice &lt;br /&gt;
&#8211;tensor-parallel-size 2 &lt;br /&gt;
&#8211;distributed-executor-backend ray &lt;br /&gt;
&#8211;gpu-memory-utilization 0.7 &lt;br /&gt;
&#8211;enable-prefix-caching &lt;br /&gt;
&#8211;load-format fastsafetensors &lt;br /&gt;
&#8211;quantization mxfp4 &lt;br /&gt;
&#8211;mxfp4-backend CUTLASS &lt;br /&gt;
&#8211;mxfp4-layers moe,qkv,o,lm_head &lt;br /&gt;
&#8211;attention-backend FLASHINFER &lt;br /&gt;
&#8211;kv-cache-dtype fp8 &lt;br /&gt;
&#8211;max-num-batched-tokens 8192 &lt;br /&gt;
&#8211;host 0.0.0.0 &lt;br /&gt;
&#8211;port 8000&lt;/p&gt;
&lt;p&gt;Runtime versions:&lt;br /&gt;
cuda: 13.1&lt;br /&gt;
nccl: (2, 29, 2)&lt;br /&gt;
python: 3.12.3&lt;br /&gt;
torch: 2.10.0a0+a36e1d39eb.nv26.01.42222806&lt;br /&gt;
vllm: 0.1.dev12774+g04f641e53.d20260720&lt;/p&gt;[6/6] Post-launch hooks — skipped</textarea></div><div class="fusion-text fusion-text-126"><p>&nbsp;</p>
<p>sparkrun successfully completed all 6 steps. The &#8216;Mode: cluster (2 nodes)&#8217; section was seen. The image was synchronized to the Worker over CX-7. <code>--tensor-parallel-size 2</code> appears in the serve command. sparkrun detected the <code>--distributed-executor-backend ray</code> flag, selected the <code>vllm-ray</code> runtime, and automatically set up the Ray cluster.</p>
<p>At TP=2, the model weights are split in two — each Spark loads ~30 GB (vs. 61 GB at TP=1). This leaves 35.97 GiB of memory for KV cache on each Spark (vs. 8.46 GiB at TP=1) — meaning more concurrent requests.</p>
<hr />
<h3>3. Monitoring the Startup Process</h3>
<p>After the model download completes, it may take a few minutes for vLLM to become ready to serve. During this process, vLLM loads the model weights into GPU memory, compiles GPU kernels, and allocates memory for inference.</p>
<p>Monitor the logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-131 > .CodeMirror, .fusion-syntax-highlighter-131 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-131 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_131" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_131" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_131" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun logs ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;tp 2</textarea></div><div class="fusion-text fusion-text-127"><p>&nbsp;</p>
<p>You will see the following in the logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-132 > .CodeMirror, .fusion-syntax-highlighter-132 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-132 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_132" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_132" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_132" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/yaml">(APIServer pid=1093) vLLM API server version 0.1.dev12774+g04f641e53.d20260720
(APIServer pid=1093) non-default args: {&#8216;quantization&#8217;: &#8216;mxfp4&#8217;, &#8216;mxfp4_backend&#8217;: &#8216;CUTLASS&#8217;,
&#8216;mxfp4_layers&#8217;: &#8216;moe,qkv,o,lm_head&#8217;, &#8216;attention_backend&#8217;: &#8216;FLASHINFER&#8217;,
&#8216;distributed_executor_backend&#8217;: &#8216;ray&#8217;, &#8216;tensor_parallel_size&#8217;: 2, &#8230;}</p>
<p>(EngineCore_DP0 pid=1407) Connecting to existing Ray cluster at address: 192.168.1.148:46379&#8230;
(EngineCore_DP0 pid=1407) Connected to Ray cluster.
(EngineCore_DP0 pid=1407) Creating a new placement group.</p>
<p>(RayWorkerWrapper pid=1510) Loading safetensors using Fastsafetensor loader: 100%|██████████| 8/8 [00:24&lt;00:00]
(RayWorkerWrapper pid=372, ip=192.168.1.163) Loading safetensors using Fastsafetensor loader: 100%|██████████| 8/8 [00:23&lt;00:00] (RayWorkerWrapper pid=372, ip=192.168.1.163) [MXFP4] lm_head quantized: torch.Size([100608, 2880]) BF16 -&gt; torch.Size([100608, 1440]) FP4 (4x smaller)
(RayWorkerWrapper pid=372, ip=192.168.1.163) Model loading took 31.99 GiB memory and 38.80 seconds</p>
<p>(RayWorkerWrapper pid=1510) torch.compile takes 116.55 s in total
(RayWorkerWrapper pid=372, ip=192.168.1.163) Available KV cache memory: 35.97 GiB
(EngineCore_DP0 pid=1407) GPU KV cache size: 2,095,472 tokens
(EngineCore_DP0 pid=1407) Maximum concurrency for 131,072 tokens per request: 30.06x</p>
<p>(APIServer pid=1093) Application startup complete.</textarea></div><div class="fusion-text fusion-text-128"><p>&nbsp;</p>
<p>If you see the <code>Application startup complete.</code> line, the model server is ready. Press <code>Ctrl+C</code> to stop watching the logs. The server will continue running in the background.</p>
<hr />
<h3>4. Testing the Model</h3>
<p>Check the container status. Containers should be running on both Sparks:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-133 > .CodeMirror, .fusion-syntax-highlighter-133 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-133 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_133" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_133" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_133" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun status</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-134 > .CodeMirror, .fusion-syntax-highlighter-134 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-134 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_134" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_134" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_134" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Job: /home/nvidia/gpt-oss-120b-eugr-mxfp4.yaml (tp=2) [631038688bb2] (2 container(s))
head 127.0.0.1 Up 4 minutes vllm-node-mxfp4
worker 192.168.1.163 Up 4 minutes vllm-node-mxfp4</p>
<p>Total: 2 container(s) across 2 host(s)</textarea></div><div class="fusion-text fusion-text-129"><p>&nbsp;</p>
<p>Both containers are running.</p>
<p>Let&#8217;s run a health check on the model:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-135 > .CodeMirror, .fusion-syntax-highlighter-135 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-135 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_135" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_135" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_135" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl -s -o /dev/null -w &#8220;HTTP %{http_code}&#8221; http://:8000/health</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-136 > .CodeMirror, .fusion-syntax-highlighter-136 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-136 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_136" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_136" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_136" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">HTTP 200</textarea></div><div class="fusion-text fusion-text-130"><p>&nbsp;</p>
<p>The model is healthy. Now let&#8217;s test the model with a prompt:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-137 > .CodeMirror, .fusion-syntax-highlighter-137 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-137 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_137" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_137" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_137" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl -s http://:8000/v1/chat/completions
-H &#8220;Content-Type: application/json&#8221;
-d &#8216;{
&#8220;model&#8221;: &#8220;openai/gpt-oss-120b&#8221;,
&#8220;messages&#8221;: [{&#8220;role&#8221;: &#8220;user&#8221;, &#8220;content&#8221;: &#8220;What is 12*17?&#8221;}],
&#8220;max_tokens&#8221;: 200
}&#8217; | python3 -m json.tool</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-138 > .CodeMirror, .fusion-syntax-highlighter-138 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-138 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_138" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_138" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_138" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="application/json">{
&#8220;id&#8221;: &#8220;chatcmpl-a2bdd35cbf300d46&#8221;,
&#8220;model&#8221;: &#8220;openai/gpt-oss-120b&#8221;,
&#8220;choices&#8221;: [
{
&#8220;index&#8221;: 0,
&#8220;message&#8221;: {
&#8220;role&#8221;: &#8220;assistant&#8221;,
&#8220;content&#8221;: &#8220;(12 times 17 = 204)&#8221;,
&#8220;reasoning&#8221;: &#8220;User asks a simple multiplication: 12*17 = 204. Provide answer.&#8221;,
&#8220;reasoning_content&#8221;: &#8220;User asks a simple multiplication: 12*17 = 204. Provide answer.&#8221;
},
&#8220;finish_reason&#8221;: &#8220;stop&#8221;
}
],
&#8220;usage&#8221;: {
&#8220;prompt_tokens&#8221;: 72,
&#8220;total_tokens&#8221;: 110,
&#8220;completion_tokens&#8221;: 38
}
}</textarea></div><div class="fusion-text fusion-text-131"><p>&nbsp;</p>
<p>The model answered <code>12 × 17 = 204</code>. Additionally, the <code>reasoning</code> and <code>reasoning_content</code> fields contain the model&#8217;s reasoning processes.</p>
<p>The API is served only through the head node (Main Spark). That is, we communicate with the model via the Main Spark. The Worker Spark only participates in the inference computation.</p>
<hr />
<h3>5. Shutdown</h3>
<p>When you are done, stop the model. Since you used <code>--tp 2</code> when launching, you need to specify <code>--tp 2</code> when stopping as well:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-139 > .CodeMirror, .fusion-syntax-highlighter-139 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-139 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_139" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_139" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_139" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun stop ~/gpt-oss-120b-eugr-mxfp4.yaml &#8211;tp 2</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-140 > .CodeMirror, .fusion-syntax-highlighter-140 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-140 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_140" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_140" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_140" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Workload stopped on 2 host(s).</textarea></div><div class="fusion-text fusion-text-132"><p>&nbsp;</p>
<p>This command stops the container and frees up memory. However, the container, the Docker image, and the model files remain on disk. This means you don&#8217;t need to re-download to start again. Simply re-run the <code>sparkrun run</code> command from step 2.</p>
<hr />
<h2>Using with Open WebUI</h2>
<p>In the previous tutorial (<a href="tutorial-vllm-openwebui-spark.md">Local LLM Serving with vLLM on DGX Spark</a>), we set up Open WebUI and connected it to vLLM&#8217;s port 8000. The GPT-OSS 120B model is also served from the same port 8000. If Open WebUI is running, it automatically detects the model and lists it as <code>openai/gpt-oss-120b</code> in the model selection menu on the chat screen. You can select the model and start using it from the browser.</p>
<hr />
<h2>Single vs Dual Spark Performance Comparison</h2>
<p>We compared the single Spark (TP=1) and dual Spark (TP=2) configurations at different concurrency levels. The measurements recorded average TTFT (Time to First Token — the time until the first token starts being generated) and TPS (Tokens Per Second — the number of tokens generated per second) values.</p>
<p>&nbsp;</p>
<table>
<thead>
<tr>
<th>Concurrency</th>
<th>Single Spark TTFT (ms)</th>
<th>Dual Spark TTFT (ms)</th>
<th>Single Spark TPS (tok/s)</th>
<th>Dual Spark TPS (tok/s)</th>
<th>TPS Improvement</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>219.57</td>
<td>167.69</td>
<td>55.65</td>
<td>69.86</td>
<td>+26%</td>
</tr>
<tr>
<td>2</td>
<td>294.41</td>
<td>227.76</td>
<td>37.05</td>
<td>51.52</td>
<td>+39%</td>
</tr>
<tr>
<td>4</td>
<td>320.61</td>
<td>255.41</td>
<td>25.20</td>
<td>37.14</td>
<td>+47%</td>
</tr>
<tr>
<td>8</td>
<td>395.31</td>
<td>291.94</td>
<td>16.88</td>
<td>26.92</td>
<td>+59%</td>
</tr>
<tr>
<td>16</td>
<td>444.10</td>
<td>319.99</td>
<td>11.67</td>
<td>19.08</td>
<td>+63%</td>
</tr>
</tbody>
</table>
<p>&nbsp;</p>
<p>The dual Spark setup provides 26% higher token throughput at concurrency 1 (single user), and as concurrency increases, this gap grows, exceeding 60% at concurrency 16, as shown.</p>
<p>The measurements were taken using the <a href="https://github.com/CordatusAI/llm-benchmark">CordatusAI LLM Benchmark Tool</a>. This tool is a benchmarking application developed by CordatusAI that tests LLM servers with OpenAI-compatible APIs. Below are screenshots of the benchmark results the application produced for our model. As can be seen, the application can test at concurrency levels from 1 to 64, and at each level it measures TTFT, ITL, TPS, latency, and throughput metrics, presenting the results as tables and graphs. It also calculates the recommended number of users the system can support. This calculation is a general estimate based on certain assumptions and does not reflect all usage scenarios. As seen in the image, the tool calculated a recommended user count of 55 for a single Spark and 120 for dual Sparks.</p>
<p><strong>Single Spark benchmark results:</strong></p>
<p>&nbsp;</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-23 hover-type-none"><img decoding="async" width="1024" height="579" title="oss_single_final_benchmark_1" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-1024x579.webp" alt class="img-responsive wp-image-1811" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-24 hover-type-none"><img decoding="async" width="1024" height="579" title="oss_single_final_benchmark_2" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-1024x579.webp" alt class="img-responsive wp-image-1812" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-133"><p>&nbsp;</p>
<p><strong>Double Spark benchmark results:</strong></p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-25 hover-type-none"><img decoding="async" width="1024" height="579" title="oss_single_final_benchmark_1" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-1024x579.webp" alt class="img-responsive wp-image-1811" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_1.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-26 hover-type-none"><img decoding="async" width="1024" height="579" title="oss_single_final_benchmark_2" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-1024x579.webp" alt class="img-responsive wp-image-1812" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/oss_single_final_benchmark_2.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div></div></div></div></div>
<p>The post <a href="https://blog.openzeka.com/en/local-gpt-oss-120b-serving-on-dgx-spark-with-sparkrun/">Local GPT-OSS 120B Serving on DGX Spark with sparkrun</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>DeepSeek-V4-Flash-0731 on 2x DGX Spark</title>
		<link>https://blog.openzeka.com/en/deepseek-v4-flash-0731-on-2x-dgx-spark/</link>
		
		<dc:creator><![CDATA[Enhar]]></dc:creator>
		<pubDate>Wed, 19 Aug 2026 13:07:16 +0000</pubDate>
				<category><![CDATA[AI Cluster]]></category>
		<category><![CDATA[Generative AI]]></category>
		<guid isPermaLink="false">https://blog.openzeka.com/en/?p=1792</guid>

					<description><![CDATA[<p>In this tutorial, you will run the DeepSeek-V4-Flas ... Continue Reading→</p>
<p>The post <a href="https://blog.openzeka.com/en/deepseek-v4-flash-0731-on-2x-dgx-spark/">DeepSeek-V4-Flash-0731 on 2x DGX Spark</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-5 fusion-flex-container nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-4 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:0px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-134"></div></div></div></div></div><div class="fusion-fullwidth fullwidth-box fusion-builder-row-6 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-5 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-135"><p>In this tutorial, you will run the DeepSeek-V4-Flash-0731 large language model across two DGX Sparks using sparkrun. This model is a Mixture-of-Experts model with 284 billion total and 13 billion active parameters. Additionally, before running it, we will examine how intelligent the model is.</p>
<p>The Spark can be used as a standalone computer by connecting a monitor and keyboard, or as a remote server accessed from another computer. In this tutorial, we will connect to the Spark remotely and install the necessary software to run the model.</p>
<p>We will use vLLM as the inference engine and sparkrun as the management tool. vLLM will load the model&#8217;s trained weights into GPU memory and expose an API that accepts external requests. sparkrun will manage Docker containers and model deployment across Sparks from the command line. This way, operations such as image synchronization, model transfer, and cluster configuration will be automated. Both will run on the Main Spark, inside Docker containers.</p>
<p>This tutorial consists of six parts:</p>
<ul>
<li><strong>Model Intelligence:</strong> Tests that measure the intelligence of AI models, how they differ from performance tests, and DeepSeek-V4-Flash-0731&#8217;s position among leading models</li>
<li><strong>Setup:</strong> Downloading the Docker image and preparing the recipe</li>
<li><strong>Running:</strong> Starting, monitoring, and testing the model on two DGX Sparks with sparkrun</li>
<li><strong>Benchmark:</strong> Performance measurement results at different concurrency levels</li>
<li><strong>Using with OpenCode:</strong> Connecting the model to a local coding assistant</li>
<li><strong>Shutdown:</strong> Stopping the services</li>
</ul>
<p>Throughout this tutorial, the primary device will be referred to as the Main Spark and the secondary device as the Worker Spark.</p>
<hr />
<h2>Model Intelligence</h2>
<h3>1. Performance Tests vs Intelligence Tests</h3>
<p>When &#8220;AI Benchmark&#8221; is mentioned, performance tests generally come to mind. In our previous tutorials, we primarily focused on these tests: we ran various models known to be intelligent on the Spark. Then we tested values such as how many tokens the model could produce per second (TPS), how long it took for the first token to arrive (TTFT), or token generation at different concurrency levels (throughput).</p>
<p>Performance tests are fundamentally about how fast a model responds. They depend on which hardware and software a particular model runs on. However, they do not tell us whether the model&#8217;s responses are correct. A model can run at 70 tok/s but produce incorrect or nonsensical responses.</p>
<p>Intelligence tests, on the other hand, are fundamentally about the accuracy of the model&#8217;s responses. They are hardware-independent — a property related to the model&#8217;s trained weights.</p>
<h3>2. Artificial Analysis and Intelligence Indices</h3>
<p><a href="https://artificialanalysis.ai">Artificial Analysis</a> is a platform that independently evaluates AI models. It is not affiliated with any model manufacturer and tests hundreds of models under the same conditions. The results are completely transparent — testing methods, question templates, and scoring rules are all publicly available. The indices they publish combine tests widely used and trusted by the academic and industrial research community. These tests come from various independent sources such as OpenAI, NYU, Stanford, and Center for AI Safety.</p>
<ul>
<li style="list-style-type: none;">
<ul>
<li><strong>Intelligence Index</strong>
<ul>
<li><strong>What It Measures:</strong> General intelligence (combination of all abilities)</li>
<li><strong>Included Tests:</strong> GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity&#8217;s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR</li>
</ul>
</li>
</ul>
</li>
</ul>
<ul>
<li><strong>Coding Index</strong>
<ul>
<li><strong>What It Measures:</strong> Coding ability</li>
<li><strong>Included Tests:</strong> Terminal-Bench v2.1, SciCode</li>
</ul>
</li>
<li><strong>Agentic Index</strong>
<ul>
<li><strong>What It Measures:</strong> Agentic ability (tool usage, planning)</li>
<li><strong>Included Tests:</strong> GDPval-AA v2, τ³-Banking</li>
</ul>
</li>
</ul>
<p>The main metric Artificial Analysis uses to evaluate model intelligence is the Intelligence Index. This index uses the weighted average of 9 separate intelligence test results. The Coding Index and Agentic Index use the tests from within the Intelligence Index that measure coding and agent usage, respectively.</p>
<p>These tests are:</p>
<p><strong>GDPval-AA v2</strong> — Real-world tasks from 44 professions. The highest-weighted component (20%).</p>
<p><strong>τ³-Banking</strong> — Multi-step agent tasks in 97 banking scenarios (14% weight).</p>
<p><strong>Terminal-Bench v2.1</strong> — 89 engineering tasks in a terminal environment (16% weight).</p>
<p><strong>SciCode</strong> — 288 coding problems across 16 science fields (8% weight).</p>
<p><strong>Humanity&#8217;s Last Exam</strong> — 2500 math, science, and social science questions prepared by experts from 500+ institutions (12% weight).</p>
<p><strong>GPQA Diamond</strong> — 198 PhD-level science questions (6% weight).</p>
<p><strong>CritPt</strong> — 71 research-level physics problems (6% weight).</p>
<p><strong>AA-Omniscience</strong> — 6000 open-ended knowledge questions (12% weight).</p>
<p><strong>AA-LCR</strong> — Information extraction and reasoning from documents of varying lengths (6% weight).</p>
<ul>
<li>Source: <a style="color: #00d114;" href="https://artificialanalysis.ai/methodology/intelligence-benchmarking"><b>Artificial Analysis Methodology</b></a></li>
</ul>
<h3>3. Intelligence Index</h3>
<p>DeepSeek-V4-Flash-0731 scored <strong>50 points</strong> on the Artificial Analysis Intelligence Index. This score ranks <strong>3rd</strong> among open-weight models.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-27 hover-type-none"><img decoding="async" width="1024" height="662" title="aa-intelligence-index" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index-1024x662.webp" alt class="img-responsive wp-image-1867" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index-200x129.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index-300x194.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index-400x259.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index-600x388.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index-768x496.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index-800x517.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index-1024x662.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index-1200x776.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-intelligence-index.webp 1205w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-136"><ul>
<li><strong>1. Kimi K3 (max)</strong>
<ul>
<li><strong>Intelligence Index:</strong> 57</li>
<li><strong>Total Params:</strong> 2800B</li>
<li><strong>Active Params:</strong> 104B</li>
</ul>
</li>
<li><strong>2. GLM-5.2 (max)</strong>
<ul>
<li><strong>Intelligence Index:</strong> 51</li>
<li><strong>Total Params:</strong> 744B</li>
<li><strong>Active Params:</strong> 40B</li>
</ul>
</li>
<li><strong>3. DeepSeek-V4-Flash-0731 (max)</strong>
<ul>
<li><strong>Intelligence Index:</strong> 50</li>
<li><strong>Total Params:</strong> 284B</li>
<li><strong>Active Params:</strong> 13B</li>
</ul>
</li>
<li><strong>4. MiniMax-M3</strong>
<ul>
<li><strong>Intelligence Index:</strong> 44</li>
<li><strong>Total Params:</strong> 428B</li>
<li><strong>Active Params:</strong> 23B</li>
</ul>
</li>
<li><strong>5. MiMo-V2.5-Pro</strong>
<ul>
<li><strong>Intelligence Index:</strong> 42</li>
<li><strong>Total Params:</strong> 1020B</li>
<li><strong>Active Params:</strong> 42B</li>
</ul>
</li>
<li><strong>6. Inkling</strong>
<ul>
<li><strong>Intelligence Index:</strong> 41</li>
<li><strong>Total Params:</strong> 975B</li>
<li><strong>Active Params:</strong> 41B</li>
</ul>
</li>
<li><strong>7. Nemotron 3 Ultra</strong>
<ul>
<li><strong>Intelligence Index:</strong> 38</li>
<li><strong>Total Params:</strong> 550B</li>
<li><strong>Active Params:</strong> 55B</li>
</ul>
</li>
<li><strong>8. Mistral Medium 3.5</strong>
<ul>
<li><strong>Intelligence Index:</strong> 30</li>
<li><strong>Total Params:</strong> 128B</li>
<li><strong>Active Params:</strong> 128B</li>
</ul>
</li>
<li><strong>9. Gemma 4 31B</strong>
<ul>
<li><strong>Intelligence Index:</strong> 29</li>
<li><strong>Total Params:</strong> 30.7B</li>
<li><strong>Active Params:</strong> 30.7B</li>
</ul>
</li>
<li><strong>10. gpt-oss-120b (high)</strong>
<ul>
<li><strong>Intelligence Index:</strong> 24</li>
<li><strong>Total Params:</strong> 117B</li>
<li><strong>Active Params:</strong> 5.1B</li>
</ul>
</li>
<li><strong>11. Command A+</strong>
<ul>
<li><strong>Intelligence Index:</strong> 23</li>
<li><strong>Total Params:</strong> 218B</li>
<li><strong>Active Params:</strong> 25B</li>
</ul>
</li>
</ul>
<p>&nbsp;</p>
<p>When we examine <strong>DeepSeek-V4-Flash-0731</strong>&#8216;s position in this ranking alongside the parameter counts of the models, an exceptional picture emerges. The two models ahead of it have much larger structures: Kimi K3, with 2.8 trillion parameters, is approximately 10 times the size of Flash and performs 8 times more computation per token with 104 billion active parameters. GLM-5.2, with 744 billion total and 40 billion active parameters, is approximately 2.6 times the size of Flash. Despite this, Flash is only 1 point behind GLM-5.2 on the general intelligence index.</p>
<p>An even more striking picture emerges when we examine the models ranked below Flash. Among the top 9 models — those with an Intelligence Index of 29 and above — Flash has the fewest active parameters. All other MoE models in the ranking use both more total and more active parameters. To find the first model with fewer active parameters than Flash, you have to go down to the 10th place, gpt-oss-120b. There, we see the model&#8217;s intelligence index is 24 — less than half of Flash&#8217;s.</p>
<p>This table has a direct implication for local inference: Parameter count is a measure of model size. Size is directly related to the memory and compute capacity required to run the model. When you choose another model from the list to achieve intelligence performance at the Intelligence Index 50 level, you need larger hardware to run it. For example, in our previous tutorials, we used 4 DGX Sparks to run GLM-5.2. Kimi K3&#8217;s 2.8 trillion parameters would require much larger hardware. DeepSeek-V4-Flash-0731, with only 284 billion total and 13 billion active parameters, is the only model at this intelligence level that can run on two DGX Sparks.</p>
<h3>4. Coding Index</h3>
<p>The Coding Index is the weighted average of the Terminal-Bench v2.1 and SciCode tests. These tests measure the model&#8217;s performance on coding problems and agent skills. DeepSeek-V4-Flash-0731 ranks <strong>2nd</strong> among open-weight models in coding.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-28 hover-type-none"><img decoding="async" width="1024" height="662" title="aa-coding-index" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index-1024x662.webp" alt class="img-responsive wp-image-1866" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index-200x129.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index-300x194.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index-400x259.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index-600x388.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index-768x496.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index-800x517.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index-1024x662.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index-1200x776.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-coding-index.webp 1205w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-137"><ul>
<li><strong>1. Kimi K3 (max)</strong>
<ul>
<li><strong>Coding Index:</strong> 76.2</li>
</ul>
</li>
<li><strong>2. DeepSeek-V4-Flash-0731 (max)</strong>
<ul>
<li><strong>Coding Index:</strong> 69.1</li>
</ul>
</li>
<li><strong>3. GLM-5.2 (max)</strong>
<ul>
<li><strong>Coding Index:</strong> 68.8</li>
</ul>
</li>
<li><strong>4. MiMo-V2.5-Pro</strong>
<ul>
<li><strong>Coding Index:</strong> 60.2</li>
</ul>
</li>
<li><strong>5. MiniMax-M3</strong>
<ul>
<li><strong>Coding Index:</strong> 58.6</li>
</ul>
</li>
<li><strong>6. Inkling</strong>
<ul>
<li><strong>Coding Index:</strong> 52.1</li>
</ul>
</li>
<li><strong>7. Nemotron 3 Ultra</strong>
<ul>
<li><strong>Coding Index:</strong> 49.3</li>
</ul>
</li>
<li><strong>8. Mistral Medium 3.5</strong>
<ul>
<li><strong>Coding Index:</strong> 46.9</li>
</ul>
</li>
<li><strong>9. Gemma 4 31B</strong>
<ul>
<li><strong>Coding Index:</strong> 43.4</li>
</ul>
</li>
<li><strong>10. gpt-oss-120b (high)</strong>
<ul>
<li><strong>Coding Index:</strong> 30.4</li>
</ul>
</li>
<li><strong>11. Command A+</strong>
<ul>
<li><strong>Coding Index:</strong> 27.8</li>
</ul>
</li>
</ul>
<p><strong>DeepSeek-V4-Flash-0731</strong> scores <strong>69.1</strong> on the Coding Index.<br />
The following <strong>GLM-5.2</strong> stays at <strong>68.8</strong> — meaning Flash surpassed<br />
even GLM-5.2, which is <strong>2.6 times its size</strong>, to claim second place.</p>
<h3>5. Agentic Index</h3>
<p>The Agentic Index is the weighted average of the <strong>GDPval-AA v2</strong> and<br />
<strong>τ³-Banking</strong> tests. It measures the model&#8217;s agentic capability.<br />
Agentic capability indicates how well the model performs as an autonomous agent —<br />
doing web research, running terminal commands, and writing code.</p>
<p>Similar to what we saw in the Coding Index,<br />
<strong>DeepSeek-V4-Flash-0731</strong> also ranks <strong>2nd</strong><br />
among open-weight models in this area, with <strong>45.7 points</strong>.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-29 hover-type-none"><img decoding="async" width="1024" height="662" title="aa-agentic-index" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index-1024x662.webp" alt class="img-responsive wp-image-1865" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index-200x129.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index-300x194.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index-400x259.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index-600x388.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index-768x496.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index-800x517.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index-1024x662.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index-1200x776.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/aa-agentic-index.webp 1205w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-138"><ul>
<li><strong>1. Kimi K3 (max)</strong>
<ul>
<li><strong>Coding Index:</strong> 76.2</li>
</ul>
</li>
<li><strong>2. DeepSeek-V4-Flash-0731 (max)</strong>
<ul>
<li><strong>Coding Index:</strong> 69.1</li>
</ul>
</li>
<li><strong>3. GLM-5.2 (max)</strong>
<ul>
<li><strong>Coding Index:</strong> 68.8</li>
</ul>
</li>
<li><strong>4. MiMo-V2.5-Pro</strong>
<ul>
<li><strong>Coding Index:</strong> 60.2</li>
</ul>
</li>
<li><strong>5. MiniMax-M3</strong>
<ul>
<li><strong>Coding Index:</strong> 58.6</li>
</ul>
</li>
<li><strong>6. Inkling</strong>
<ul>
<li><strong>Coding Index:</strong> 52.1</li>
</ul>
</li>
<li><strong>7. Nemotron 3 Ultra</strong>
<ul>
<li><strong>Coding Index:</strong> 49.3</li>
</ul>
</li>
<li><strong>8. Mistral Medium 3.5</strong>
<ul>
<li><strong>Coding Index:</strong> 46.9</li>
</ul>
</li>
<li><strong>9. Gemma 4 31B</strong>
<ul>
<li><strong>Coding Index:</strong> 43.4</li>
</ul>
</li>
<li><strong>10. gpt-oss-120b (high)</strong>
<ul>
<li><strong>Coding Index:</strong> 30.4</li>
</ul>
</li>
<li><strong>11. Command A+</strong>
<ul>
<li><strong>Coding Index:</strong> 27.8</li>
</ul>
</li>
</ul>
<p><strong>DeepSeek-V4-Flash-0731</strong> scores <strong>69.1</strong> on the Coding Index.<br />
The following <strong>GLM-5.2</strong> stays at <strong>68.8</strong> — meaning Flash surpassed<br />
even GLM-5.2, which is <strong>2.6 times its size</strong>, to claim second place.</p>
<h3>5. Agentic Index</h3>
<p>The Agentic Index is the weighted average of the <strong>GDPval-AA v2</strong> and<br />
<strong>τ³-Banking</strong> tests. It measures the model&#8217;s agentic capability.</p>
<p>Agentic capability indicates how well the model performs as an autonomous agent —<br />
doing web research, running terminal commands, and writing code.<br />
Similar to what we saw in the Coding Index, <strong>DeepSeek-V4-Flash-0731</strong><br />
also ranks <strong>2nd</strong> among open-weight models in this area, with<br />
<strong>45.7 points</strong>.</p>
</div><div class="fusion-text fusion-text-139"><h2>Setup</h2>
<h3>Prerequisites</h3>
<p>Previous tutorials covered the processes of connecting to the Spark, installing sparkrun, and configuring a multi-Spark cluster step by step. In this tutorial, we assume all these steps have been completed and your setup is ready.</p>
<p>This tutorial requires 2 DGX Sparks. We will refer to the primary device as the and the other device as . Note the IP address of each device beforehand.</p>
<p>Downloading the Docker Image</p>
<p>DeepSeek-V4-Flash-0731&#8217;s MoE architecture and MLA (Multi-Head Latent Attention) structure require kernels compiled for GB10 (SM 12.1). Standard vLLM images do not include these kernels. Therefore, we will use a custom Docker image that includes the required B12X kernels.</p>
<p>Run the following command to pull the Docker image prepared by OpenZeka:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-141 > .CodeMirror, .fusion-syntax-highlighter-141 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-141 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_141" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_141" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_141" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker pull registry.cordata.ai/spark-cluster/vllm-node-b12x:latest</textarea></div><div class="fusion-text fusion-text-140"><p>&nbsp;</p>
<p>Verify the image:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-142 > .CodeMirror, .fusion-syntax-highlighter-142 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-142 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_142" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_142" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_142" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker images registry.cordata.ai/spark-cluster/vllm-node-b12x:latest</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-143 > .CodeMirror, .fusion-syntax-highlighter-143 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-143 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_143" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_143" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_143" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">REPOSITORY TAG IMAGE ID CREATED SIZE
registry.cordata.ai/spark-cluster/vllm-node-b12x latest 91ddbb01c9d6 2 minutes ago 34.5GB</textarea></div><div class="fusion-text fusion-text-141"><h3>2. Preparing the Recipe</h3>
<p>A recipe is a YAML file that defines how sparkrun will run the model. The model, Docker image, vLLM flags, and memory settings are all consolidated in a single file. For DeepSeek-V4-Flash-0731, this recipe uses tensor parallelism across 2 nodes (two Sparks share the model weight matrices), runs the model with the maximum context window, and accelerates model inference with DSpark speculative decoding (k=5).</p>
<p>Save the recipe file using the following command:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-144 > .CodeMirror, .fusion-syntax-highlighter-144 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-144 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_144" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_144" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_144" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/yaml">cat &amp;gt; ~/deepseek-v4-flash-0731-eugr-b12x.yaml &amp;lt;&amp;lt; &#8216;RECIPE'&lt;br /&gt;
# DeepSeek V4 Flash 0731 B12X TP2&lt;br /&gt;
# Usage:&lt;br /&gt;
# sparkrun run ~/deepseek-v4-flash-0731-eugr-b12x.yaml # 2-node cluster (TP=2)&lt;br /&gt;
recipe_version: &#8220;2&#8221;&lt;br /&gt;
model: deepseek-ai/DeepSeek-V4-Flash-0731&lt;br /&gt;
runtime: vllm-distributed&lt;br /&gt;
min_nodes: 2&lt;br /&gt;
max_nodes: 2&lt;br /&gt;
container: registry.cordata.ai/spark-cluster/vllm-node-b12x:latest&lt;/p&gt;
&lt;p&gt;metadata:&lt;br /&gt;
description: vLLM serving deepseek-ai/DeepSeek-V4-Flash-0731 on a dual Sparks using B12X docker&lt;br /&gt;
model_params: 284B&lt;br /&gt;
model_dtype: nvfp4&lt;br /&gt;
kv_dtype: fp8&lt;/p&gt;
&lt;p&gt;defaults:&lt;br /&gt;
port: 8000&lt;br /&gt;
host: 0.0.0.0&lt;br /&gt;
tensor_parallel: 2&lt;br /&gt;
gpu_memory_utilization: 0.85&lt;br /&gt;
max_model_len: auto&lt;br /&gt;
block_size: 256&lt;br /&gt;
max_num_seqs: 8&lt;br /&gt;
max_num_batched_tokens: 8192&lt;br /&gt;
max_cudagraph_capture_size: 64&lt;br /&gt;
num_speculative_tokens: 5&lt;br /&gt;
reasoning_config: &#8216;{&#8220;reasoning_parser&#8221;:&#8221;deepseek_v4&#8243;,&#8221;reasoning_start_str&#8221;:&#8221;&#8221;,&#8221;reasoning_end_str&#8221;:&#8221;&#8221;}'&lt;br /&gt;
compilation_config: &#8216;{&#8220;cudagraph_mode&#8221;:&#8221;FULL_AND_PIECEWISE&#8221;,&#8221;custom_ops&#8221;:[&#8220;all&#8221;]}'&lt;br /&gt;
speculative_config: &#8216;{&#8220;method&#8221;:&#8221;dspark&#8221;,&#8221;num_speculative_tokens&#8221;:5,&#8221;draft_sample_method&#8221;:&#8221;probabilistic&#8221;,&#8221;attention_backend&#8221;:&#8221;B12X_MLA_SPARSE&#8221;}'&lt;/p&gt;
&lt;p&gt;env:&lt;br /&gt;
CUTE_DSL_ARCH: &#8220;sm_121a&#8221;&lt;br /&gt;
VLLM_USE_AOT_COMPILE: &#8220;1&#8221;&lt;br /&gt;
VLLM_USE_BREAKABLE_CUDAGRAPH: &#8220;0&#8221;&lt;br /&gt;
VLLM_USE_MEGA_AOT_ARTIFACT: &#8220;-1&#8243;&lt;br /&gt;
VLLM_MEMORY_PROFILE_INCLUDE_ATTN: &#8220;1&#8221;&lt;br /&gt;
VLLM_USE_FLASHINFER_SAMPLER: &#8220;1&#8221;&lt;br /&gt;
VLLM_USE_B12X_WO_PROJECTION: &#8220;1&#8221;&lt;br /&gt;
VLLM_USE_B12X_MHC: &#8220;1&#8221;&lt;br /&gt;
VLLM_USE_B12X_FP8_GEMM: &#8220;1&#8221;&lt;br /&gt;
VLLM_USE_B12X_MOE: &#8220;1&#8221;&lt;br /&gt;
VLLM_USE_B12X_SPARSE_INDEXER: &#8220;1&#8221;&lt;br /&gt;
VLLM_USE_V2_MODEL_RUNNER: &#8220;1&#8221;&lt;br /&gt;
B12X_MLA_SM120_UNIFIED: &#8220;1&#8221;&lt;br /&gt;
B12X_MOE_FORCE_A8: &#8220;1&#8221;&lt;br /&gt;
VLLM_WORKER_MULTIPROC_METHOD: &#8220;spawn&#8221;&lt;/p&gt;
&lt;p&gt;command: |&lt;br /&gt;
vllm serve {model}&lt;br /&gt;
&#8211;served-model-name {model}&lt;br /&gt;
&#8211;host {host}&lt;br /&gt;
&#8211;port {port}&lt;br /&gt;
&#8211;trust-remote-code&lt;br /&gt;
&#8211;tensor-parallel-size {tensor_parallel}&lt;br /&gt;
&#8211;kv-cache-dtype fp8&lt;br /&gt;
&#8211;block-size {block_size}&lt;br /&gt;
&#8211;max-model-len {max_model_len}&lt;br /&gt;
&#8211;max-num-seqs {max_num_seqs}&lt;br /&gt;
&#8211;max-num-batched-tokens {max_num_batched_tokens}&lt;br /&gt;
&#8211;gpu-memory-utilization {gpu_memory_utilization}&lt;br /&gt;
&#8211;enable-prefix-caching&lt;br /&gt;
&#8211;tokenizer-mode deepseek_v4&lt;br /&gt;
&#8211;tool-call-parser deepseek_v4&lt;br /&gt;
&#8211;enable-auto-tool-choice&lt;br /&gt;
&#8211;reasoning-parser deepseek_v4&lt;br /&gt;
&#8211;reasoning-config &#8216;{reasoning_config}'&lt;br /&gt;
&#8211;default-chat-template-kwargs.thinking=true&lt;br /&gt;
&#8211;default-chat-template-kwargs.reasoning_effort=high&lt;br /&gt;
&#8211;load-format instanttensor&lt;br /&gt;
&#8211;moe-backend b12x&lt;br /&gt;
&#8211;linear-backend b12x&lt;br /&gt;
&#8211;attention-backend B12X_MLA_SPARSE&lt;br /&gt;
&#8211;max-cudagraph-capture-size {max_cudagraph_capture_size}&lt;br /&gt;
&#8211;compilation-config &#8216;{compilation_config}'&lt;br /&gt;
&#8211;speculative-config &#8216;{speculative_config}'&lt;br /&gt;
RECIPE</textarea></div><div class="fusion-text fusion-text-142"><p>&nbsp;</p>
<p>Verify the file:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-145 > .CodeMirror, .fusion-syntax-highlighter-145 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-145 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_145" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_145" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_145" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cat ~/deepseek-v4-flash-0731-eugr-b12x.yaml</textarea></div><div class="fusion-text fusion-text-143"><p>You will see the recipe file you saved.</p>
<hr />
<h2>Running</h2>
<h3>1. Clearing the Cache</h3>
<p>Before starting vLLM, flush the filesystem cache on each Spark. The main reason for doing this is the DGX Spark&#8217;s unified memory architecture (UMA): the operating system caches model files read from disk in RAM. vLLM loads the model weights from here into GPU memory. After loading is complete, the cached data remains in RAM even though it will not be used again. On systems with separate memory, this is not significant — since inference runs in GPU memory, RAM utilization does not affect performance. On the Spark, however, the CPU and GPU share the same RAM, so the cache reduces the memory available to the GPU. This command frees the cache, providing maximum memory for vLLM.</p>
<p>Clear the cache on the Main Spark:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-146 > .CodeMirror, .fusion-syntax-highlighter-146 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-146 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_146" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_146" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_146" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sudo sh -c &#8216;sync; echo 3 &gt; /proc/sys/vm/drop_caches&#8217;</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-147 > .CodeMirror, .fusion-syntax-highlighter-147 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-147 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_147" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_147" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_147" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">[sudo] password for nvidia:</textarea></div><div class="fusion-text fusion-text-144"><p>&nbsp;</p>
<p>From the Main Spark, apply the same cleanup on the Worker Spark via SSH:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-148 > .CodeMirror, .fusion-syntax-highlighter-148 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-148 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_148" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_148" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_148" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">ssh nvidia@ &#8220;sudo sh -c &#8216;sync; echo 3 &gt; /proc/sys/vm/drop_caches'&#8221;</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-149 > .CodeMirror, .fusion-syntax-highlighter-149 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-149 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_149" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_149" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_149" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">[sudo] password for nvidia:</textarea></div><div class="fusion-text fusion-text-145"><hr />
<h3>2. Pre-launch Checks</h3>
<p>Verify that sparkrun correctly parsed the recipe and that the memory budget is suitable:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-150 > .CodeMirror, .fusion-syntax-highlighter-150 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-150 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_150" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_150" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_150" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun show ~/deepseek-v4-flash-0731-eugr-b12x.yaml</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-151 > .CodeMirror, .fusion-syntax-highlighter-151 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-151 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_151" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_151" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_151" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Name: /home/nvidia/deepseek-v4-flash-0731-eugr-b12x.yaml
Description: vLLM serving deepseek-ai/DeepSeek-V4-Flash-0731 on a dual Sparks using B12X docker
Runtime: vllm-distributed
Model: deepseek-ai/DeepSeek-V4-Flash-0731
Container: registry.cordata.ai/spark-cluster/vllm-node-b12x:latest
Nodes: 2 &#8211; 2</p>
<p>&#8230;..
&#8230;..
&#8230;..</p>
<p>VRAM Estimation:
Model dtype: nvfp4
Model params: 284,000,000,000
KV cache dtype: fp8
Architecture: 43 layers, 1 KV heads, 512 head_dim
Model weights: 132.25 GB
Tensor parallel: 2
Per-GPU total: 66.12 GB
DGX Spark fit: YES</p>
<p>GPU Memory Budget:
gpu_memory_utilization: 85%
Usable GPU memory: 102.8 GB (121 GB x 85%)
Available for KV: 36.7 GB
Max context tokens: 1,791,167</textarea></div><div class="fusion-text fusion-text-146"><p>&nbsp;</p>
<p>sparkrun correctly parsed the recipe, selected vllm-distributed, and confirmed &#8216;DGX Spark fit: YES&#8217;.</p>
<hr />
<h3>3. Starting the Model</h3>
<p>Now let&#8217;s start the model:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-152 > .CodeMirror, .fusion-syntax-highlighter-152 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-152 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_152" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_152" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_152" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun run ~/deepseek-v4-flash-0731-eugr-b12x.yaml &#8211;no-follow</textarea></div><div class="fusion-text fusion-text-147"><p>&nbsp;</p>
<p>sparkrun automatically synchronizes the image to the Worker (skips if same ID), downloads the model to the head node and distributes it to the Worker (skips if already present), configures NCCL for CX-7 interfaces, and launches containers on each Spark.</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-153 > .CodeMirror, .fusion-syntax-highlighter-153 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-153 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_153" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_153" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_153" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/yaml">Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.&lt;br /&gt;
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.&lt;br /&gt;
sparkrun v0.3.1&lt;/p&gt;
&lt;p&gt;Runtime: vllm-distributed&lt;br /&gt;
Image: registry.cordata.ai/spark-cluster/vllm-node-b12x:latest&lt;br /&gt;
Model: deepseek-ai/DeepSeek-V4-Flash-0731&lt;br /&gt;
Mode: cluster (2 nodes)&lt;br /&gt;
Platform: DGX Spark (NVIDIA GB10, NCCL)&lt;br /&gt;
Scheduler: occupancy-sparse&lt;/p&gt;
&lt;p&gt;VRAM Estimation:&lt;br /&gt;
Model dtype: nvfp4&lt;br /&gt;
Model params: 284,000,000,000&lt;br /&gt;
KV cache dtype: fp8&lt;br /&gt;
Architecture: 43 layers, 1 KV heads, 512 head_dim&lt;br /&gt;
Model weights: 132.25 GB&lt;br /&gt;
Tensor parallel: 2&lt;br /&gt;
Per-GPU total: 66.12 GB&lt;br /&gt;
DGX Spark fit: YES&lt;/p&gt;
&lt;p&gt;GPU Memory Budget:&lt;br /&gt;
gpu_memory_utilization: 85%&lt;br /&gt;
Usable GPU memory: 102.8 GB (121 GB x 85%)&lt;br /&gt;
Available for KV: 36.7 GB&lt;br /&gt;
Max context tokens: 1,791,167&lt;/p&gt;
&lt;p&gt;Per-host fit:&lt;br /&gt;
: ranks=1, per-rank=66.1 GB, accelerator=121.0 GB @85% -&amp;gt; usable=102.8 GB, headroom=36.7 GB [OK]
: ranks=1, per-rank=66.1 GB, accelerator=121.0 GB @85% -&amp;gt; usable=102.8 GB, headroom=36.7 GB [OK]
&lt;p&gt;Hosts: default cluster &#8216;default'&lt;br /&gt;
Head:&lt;br /&gt;
Workers:&lt;/p&gt;[1/6] Preparing&lt;br /&gt;
done (0.0s)[2/6] Building — skipped (no builder)[3/6] Distributing resources&lt;br /&gt;
Distributing image registry.cordata.ai/spark-cluster/vllm-node-b12x:latest to 2 host(s)&lt;br /&gt;
Container image stale on 1 of 2 host(s), syncing&lt;br /&gt;
Distributing model deepseek-ai/DeepSeek-V4-Flash-0731 to 2 host(s)&lt;br /&gt;
Fetching 74 files: 100%|██████████| 74/74 [00:00&amp;lt;00:00, 5336.17it/s]
Model synced to 2 host(s)&lt;br /&gt;
done (185.3s)[4/6] Syncing tuning configs&lt;br /&gt;
done (0.0s)[5/6] Launching vllm runtime&lt;br /&gt;
Step 1/7: Cleaning up existing containers&lt;br /&gt;
Step 2/7: Detecting InfiniBand&lt;br /&gt;
Step 3/7: Detecting head node IP&lt;br /&gt;
Step 4/7: Launching containers&lt;br /&gt;
Step 5/7: Running pre-serve hooks&lt;br /&gt;
Step 6/7: Starting head node serve&lt;br /&gt;
Step 7/7: Starting worker nodes&lt;br /&gt;
done (69.9s)&lt;br /&gt;
Cluster: sparkrun_81fb38c99696b8c8_59ab24d3b0b4&lt;/p&gt;
&lt;p&gt;Serve command:&lt;br /&gt;
vllm serve deepseek-ai/DeepSeek-V4-Flash-0731&lt;br /&gt;
&#8211;served-model-name deepseek-ai/DeepSeek-V4-Flash-0731&lt;br /&gt;
&#8211;host 0.0.0.0&lt;br /&gt;
&#8211;port 8000&lt;br /&gt;
&#8211;trust-remote-code&lt;br /&gt;
&#8211;tensor-parallel-size 2&lt;br /&gt;
&#8211;kv-cache-dtype fp8&lt;br /&gt;
&#8211;block-size 256&lt;br /&gt;
&#8211;max-model-len auto&lt;br /&gt;
&#8211;max-num-seqs 8&lt;br /&gt;
&#8211;max-num-batched-tokens 8192&lt;br /&gt;
&#8211;gpu-memory-utilization 0.85&lt;br /&gt;
&#8211;enable-prefix-caching&lt;br /&gt;
&#8211;tokenizer-mode deepseek_v4&lt;br /&gt;
&#8211;tool-call-parser deepseek_v4&lt;br /&gt;
&#8211;enable-auto-tool-choice&lt;br /&gt;
&#8211;reasoning-parser deepseek_v4&lt;br /&gt;
&#8211;reasoning-config &#8216;{&#8220;reasoning_parser&#8221;:&#8221;deepseek_v4&#8243;,&#8221;reasoning_start_str&#8221;:&#8221;&#8221;,&#8221;reasoning_end_str&#8221;:&#8221;&#8221;}'&lt;br /&gt;
&#8211;default-chat-template-kwargs.thinking=true&lt;br /&gt;
&#8211;default-chat-template-kwargs.reasoning_effort=high&lt;br /&gt;
&#8211;load-format instanttensor&lt;br /&gt;
&#8211;moe-backend b12x&lt;br /&gt;
&#8211;linear-backend b12x&lt;br /&gt;
&#8211;attention-backend B12X_MLA_SPARSE&lt;br /&gt;
&#8211;max-cudagraph-capture-size 64&lt;br /&gt;
&#8211;compilation-config &#8216;{&#8220;cudagraph_mode&#8221;:&#8221;FULL_AND_PIECEWISE&#8221;,&#8221;custom_ops&#8221;:[&#8220;all&#8221;]}'&lt;br /&gt;
&#8211;speculative-config &#8216;{&#8220;method&#8221;:&#8221;dspark&#8221;,&#8221;num_speculative_tokens&#8221;:5,&#8221;draft_sample_method&#8221;:&#8221;probabilistic&#8221;,&#8221;attention_backend&#8221;:&#8221;B12X_MLA_SPARSE&#8221;}'&lt;/p&gt;
&lt;p&gt;Runtime versions:&lt;br /&gt;
cuda: 13.0&lt;br /&gt;
nccl: (2, 29, 7)&lt;br /&gt;
python: 3.12.3&lt;br /&gt;
torch: 2.12.0+cu130&lt;br /&gt;
vllm: 0.1.dev19023+g30038602b.d20260804&lt;/p&gt;[6/6] Post-launch hooks — skipped</textarea></div><div class="fusion-text fusion-text-148"><p>&nbsp;</p>
<p>sparkrun successfully completed all 6 steps. &#8216;Mode: cluster (2 nodes)&#8217; was observed. sparkrun synchronized the image to the Worker over the CX-7 network. All flags were correctly resolved within the serve command.</p>
<p>After the model download is complete, it may take a few minutes for vLLM to become ready for serving. During this process, vLLM loads the model weights into GPU memory, compiles GPU kernels, and allocates memory for inference.</p>
<p>To watch vLLM logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-154 > .CodeMirror, .fusion-syntax-highlighter-154 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-154 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_154" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_154" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_154" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun logs ~/deepseek-v4-flash-0731-eugr-b12x.yaml</textarea></div><div class="fusion-text fusion-text-149"><p>&nbsp;</p>
<p>You will see the following in the logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-155 > .CodeMirror, .fusion-syntax-highlighter-155 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-155 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_155" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_155" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_155" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">(APIServer pid=89) INFO 08-05 12:58:36 [api_utils.py:345]
(APIServer pid=89) INFO 08-05 12:58:36 [api_utils.py:345] █ █ █▄ ▄█
(APIServer pid=89) INFO 08-05 12:58:36 [api_utils.py:345] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.1.dev19023+g30038602b.d20260804
(APIServer pid=89) INFO 08-05 12:58:36 [api_utils.py:345] █▄█▀ █ █ █ █ model deepseek-ai/DeepSeek-V4-Flash-0731
(APIServer pid=89) INFO 08-05 12:58:36 [api_utils.py:345] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀
(APIServer pid=89) INFO 08-05 12:58:36 [api_utils.py:345]
<p>(APIServer pid=89) INFO 08-05 12:58:47 [model.py:622] Resolved architecture: DeepseekV4ForCausalLM
(APIServer pid=89) INFO 08-05 12:58:47 [model.py:1794] Using max model len 1048576
(APIServer pid=89) INFO 08-05 12:58:47 [cache.py:286] Using fp8 data type to store kv cache. It reduces the GPU memory footprint and boosts the performance. Meanwhile, it may cause accuracy drop without a proper scaling factor</p>
<p>(APIServer pid=89) INFO 08-05 12:58:56 [model.py:622] Resolved architecture: DeepseekV4MTPModel
(APIServer pid=89) INFO 08-05 12:58:56 [model.py:1794] Using max model len 1048576
(APIServer pid=89) INFO 08-05 12:59:06 [scheduler.py:252] Chunked prefill is enabled with max_num_batched_tokens=8192.
(APIServer pid=89) INFO 08-05 12:59:06 [vllm.py:1162] Asynchronous scheduling is enabled.</p>
<p>&#8230;..
&#8230;..
&#8230;..</p>
<p>(Worker_TP0 pid=321) Capturing CUDA graphs (PIECEWISE): 100%|██████████| 11/11 [00:04&lt;00:00, 2.25it/s]
(Worker_TP0 pid=321) Capturing CUDA graphs (FULL): 100%|██████████| 8/8 [00:11&lt;00:00, 1.38s/it]
(Worker_TP0 pid=321) Capturing dspark CUDA graphs (FULL): 100%|██████████| 8/8 [00:00&lt;00:00, 17.14it/s]
(Worker_TP0 pid=321) INFO 08-05 13:03:40 [model_runner.py:1067] Graph capturing finished in 17 secs, took 0.07 GiB</p>
<p>(EngineCore pid=284) INFO 08-05 13:03:48 [core.py:345] init engine (profile, create kv cache, warmup model) took 154.21 s (compilation: 16.84 s)</p>
<p>(APIServer pid=89) INFO 08-05 13:03:49 [parser_manager.py:37] &#8220;auto&#8221; tool choice has been enabled.
(APIServer pid=89) INFO 08-05 13:03:49 [api_server.py:684] Starting vLLM server on http://0.0.0.0:8000
(APIServer pid=89) INFO: Started server process [89]
(APIServer pid=89) INFO: Waiting for application startup.
(APIServer pid=89) INFO: Application startup complete.</textarea></div><div class="fusion-text fusion-text-150"><p>&nbsp;</p>
<p>Once you see the <code>Application startup complete.</code> line, the model server is ready.</p>
<p>Finally, verify that the server is running:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-156 > .CodeMirror, .fusion-syntax-highlighter-156 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-156 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_156" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_156" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_156" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl -s -o /dev/null -w &#8220;HTTP %{http_code}&#8221; http://:8000/health</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-157 > .CodeMirror, .fusion-syntax-highlighter-157 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-157 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_157" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_157" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_157" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">HTTP 200</textarea></div><div class="fusion-text fusion-text-151"><p>&nbsp;</p>
<p>The &#8216;HTTP 200&#8217; response indicates that the server is healthy.</p>
<p>DeepSeek-V4-Flash-0731 is now running and serving from port 8000 on the Main Spark.</p>
<hr />
<h2>Benchmark</h2>
<p>We tested the DeepSeek-V4-Flash-0731 model running on two Sparks at different concurrency levels. The measurements recorded average TTFT (Time to First Token — time to start producing the first token) and TPS (Tokens Per Second — number of tokens produced per second) values.</p>
<p>&nbsp;</p>
<table>
<thead>
<tr>
<th>Concurrency</th>
<th>Avg TTFT (ms)</th>
<th>Avg TPS (tok/s)</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>308</td>
<td>47.3</td>
</tr>
<tr>
<td>2</td>
<td>416</td>
<td>32.4</td>
</tr>
<tr>
<td>4</td>
<td>532</td>
<td>22.5</td>
</tr>
<tr>
<td>8</td>
<td>686</td>
<td>15.7</td>
</tr>
</tbody>
</table>
<p>&nbsp;</p>
<p>The measurements were taken using the <a href="https://github.com/CordatusAI/llm-benchmark">CordatusAI LLM Benchmark Tool</a>. This tool is a benchmarking application developed by CordatusAI that tests LLM servers with OpenAI-compatible APIs. As can be seen, at single concurrency, the value exceeded 47 tokens per second.</p>
<hr />
<h2>Using with OpenCode</h2>
<p><a href="https://opencode.ai">OpenCode</a> is an open-source AI coding assistant that runs from the terminal. It has features such as writing code, editing files, running terminal commands, and analyzing codebases. Unlike cloud services, OpenCode runs on your own computer and you choose which LLM it works with.</p>
<p>If you want to take advantage of the coding and agentic capabilities of the DeepSeek-V4-Flash-0731 model, you can connect OpenCode to this model to get a local coding assistant.</p>
<h3>1. Installation and Configuration</h3>
<p>Install OpenCode on your computer:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-158 > .CodeMirror, .fusion-syntax-highlighter-158 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-158 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_158" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_158" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_158" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">curl -fsSL https://opencode.ai/install | bash</textarea></div><div class="fusion-text fusion-text-152"><p>&nbsp;</p>
<p>Verify the installation:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-159 > .CodeMirror, .fusion-syntax-highlighter-159 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-159 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_159" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_159" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_159" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">opencode &#8211;version</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-160 > .CodeMirror, .fusion-syntax-highlighter-160 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-160 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_160" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_160" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_160" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">1.18.14</textarea></div><div class="fusion-text fusion-text-153"><p>&nbsp;</p>
<p>Save the following JSON to <code>~/.config/opencode/opencode.json</code>. Create the file if it doesn&#8217;t exist. Replace <code></code> with your Spark&#8217;s IP address:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-161 > .CodeMirror, .fusion-syntax-highlighter-161 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-161 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_161" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_161" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_161" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="application/json">{
&#8220;$schema&#8221;: &#8220;https://opencode.ai/config.json&#8221;,
&#8220;model&#8221;: &#8220;local/deepseek-v4-flash&#8221;,
&#8220;provider&#8221;: {
&#8220;local&#8221;: {
&#8220;npm&#8221;: &#8220;@ai-sdk/openai-compatible&#8221;,
&#8220;name&#8221;: &#8220;DGX Spark&#8221;,
&#8220;options&#8221;: {
&#8220;baseURL&#8221;: &#8220;http://:8000/v1&#8221;
},
&#8220;models&#8221;: {
&#8220;deepseek-v4-flash&#8221;: {
&#8220;id&#8221;: &#8220;deepseek-ai/DeepSeek-V4-Flash-0731&#8221;,
&#8220;name&#8221;: &#8220;DeepSeek-V4-Flash-0731&#8221;,
&#8220;reasoning&#8221;: true,
&#8220;tool_call&#8221;: true,
&#8220;limit&#8221;: {
&#8220;context&#8221;: 1048576,
&#8220;output&#8221;: 32768
}
}
}
}
}
}</textarea></div><div class="fusion-text fusion-text-154"><blockquote>
<p><strong>Important:</strong> If you already have a <code>~/.config/opencode/opencode.json</code> file on your computer, back it up to avoid losing it: <code>cp ~/.config/opencode/opencode.json ~/.config/opencode/opencode.json.bak</code></p>
</blockquote>
<h3>2. Usage</h3>
<p>Start OpenCode from the project directory you want to work in:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-162 > .CodeMirror, .fusion-syntax-highlighter-162 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-162 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_162" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_162" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_162" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">cd ~/my-project
opencode</textarea></div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-30 hover-type-none"><img decoding="async" width="1024" height="579" title="opencode-baslangic-ekrani" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-1024x579.webp" alt class="img-responsive wp-image-1869" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-baslangic-ekrani.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-155"><p>&nbsp;</p>
<p>Type your message in the input box and press <code>Enter</code> to start chatting with the model:</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-31 hover-type-none"><img decoding="async" width="1024" height="579" title="opencode-ilk-sohbet" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-1024x579.webp" alt class="img-responsive wp-image-1873" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-ilk-sohbet.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-156"><p>&nbsp;</p>
<p>In OpenCode, you switch between <strong>Build</strong> and <strong>Plan</strong> modes with the <code>Tab</code> key. Start complex tasks in Plan mode. In this mode, OpenCode only has permissions for analysis, reading, research, and planning. Once you approve the plan, you can switch to Build mode. In this mode, file creation, editing, and command execution features are also enabled.</p>
<p>You can use OpenCode commands with <code>/</code>. Use <code>/models</code> to see available models, <code>/sessions</code> to switch to past sessions, and <code>/undo</code> to undo changes.</p>
<p>Use <code>/help</code> for other commands:</p>
<table>
<thead>
<tr>
<th>Command</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td><code><strong>/models</strong></code></td>
<td>Lists available models, switches model</td>
</tr>
<tr>
<td><strong><code>/init</code></strong></td>
<td>Analyzes project, creates <code>AGENTS.md</code></td>
</tr>
<tr>
<td><code><strong>/compact</strong></code></td>
<td>Compresses context</td>
</tr>
<tr>
<td><strong><code>/share</code></strong></td>
<td>Converts session to shareable link</td>
</tr>
<tr>
<td><strong><code>/export</code></strong></td>
<td>Exports session as Markdown</td>
</tr>
<tr>
<td><strong><code>/thinking</code></strong></td>
<td>Show/hide model&#8217;s chain of thought</td>
</tr>
</tbody>
</table>
<p>&nbsp;</p>
<p>You can also reference files with <code>@</code> in your messages and run terminal commands directly with <code>!</code>.</p>
<p>When you give a complex task, OpenCode first analyzes it:</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-32 hover-type-none"><img decoding="async" width="1024" height="579" title="opencode-gorev-isleniyor" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-1024x579.webp" alt class="img-responsive wp-image-1870" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-gorev-isleniyor.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-157"><p>&nbsp;</p>
<p>Then it presents an implementation plan and asks you questions about ambiguous points:</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-33 hover-type-none"><img decoding="async" width="1024" height="579" title="opencode-plan-ve-sorular" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-1024x579.webp" alt class="img-responsive wp-image-1874" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-plan-ve-sorular.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-158"><p>&nbsp;</p>
<p>After answering the questions, you can switch to Build mode and allow file creation. After the model creates the files, you can verify the project:</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-34 hover-type-none"><img decoding="async" width="1024" height="579" title="opencode-proje-tamamlandi" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-1024x579.webp" alt class="img-responsive wp-image-1875" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-proje-tamamlandi.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-159"><hr />
<h3>3. Example Usage</h3>
<p>Now that you&#8217;ve learned OpenCode, let&#8217;s give a more challenging task to test the model&#8217;s coding and intelligence capabilities. Start a new session, switch to Build mode with <code>Tab</code>, and send the following prompt:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-163 > .CodeMirror, .fusion-syntax-highlighter-163 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-163 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_163" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_163" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_163" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Write a complete, single-file HTML page, named hexagon.html, containing a JavaScript Canvas simulation of a single red ball with radius 20 pixels bouncing inside a spinning blue hexagon against a solid black background and no other objects in the scene. The hexagon should be perfectly centered, have a radius of 250 pixels, and rotate clockwise at a constant speed of 60 degrees per second. Its motion should follow comprehensive and realistic 2D physics such as mass, earth&#8217;s gravity, friction, restitution, and momentum. Construct and implement accurate vector-based physics, paying particular attention to collisions between the ball and the moving hexagon boundaries: when the ball strikes a rotating boundary, the response must account for the wall&#8217;s instantaneous velocity at the contact point so that the ball correctly inherits the appropriate tangential momentum and velocity from the hexagon&#8217;s rotation. The simulation should run smoothly at approximately 60 FPS using requestAnimationFrame, and all HTML, CSS, and Vanilla JavaScript must be fully self-contained in a single file with zero external dependencies.</textarea></div><div class="fusion-text fusion-text-160"><p>This prompt requires the model to correctly integrate multiple challenging topics within a single file, including vector-based physics, collision calculation with rotating objects, and Canvas API usage.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-35 hover-type-none"><img decoding="async" width="1024" height="579" title="opencode-hexagon-prompt" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-1024x579.webp" alt class="img-responsive wp-image-1872" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-prompt.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-161"><p>&nbsp;</p>
<p>The model analyzes the prompt and starts creating the file:</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-36 hover-type-none"><img decoding="async" width="1024" height="579" title="opencode-hexagon-created" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-1024x579.webp" alt class="img-responsive wp-image-1871" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/opencode-hexagon-created.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-162"><p>&nbsp;</p>
<p>You can test the simulation by opening the generated <strong><code>hexagon.html</code></strong> file in your browser. You will see the ball bouncing inside the rotating hexagon following realistic physics rules:</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-37 hover-type-none"><img decoding="async" width="1024" height="579" title="hexagon" src="https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-1024x579.webp" alt class="img-responsive wp-image-1868" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/08/hexagon.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-163"><p>&nbsp;</p>
<p>You can also try the same prompt with other models and compare the results.</p>
<hr />
<h2>Shutdown</h2>
<p>When you are done, stop the model:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-164 > .CodeMirror, .fusion-syntax-highlighter-164 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-164 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_164" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_164" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_164" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun stop ~/deepseek-v4-flash-0731-eugr-b12x.yaml</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-165 > .CodeMirror, .fusion-syntax-highlighter-165 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-165 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_165" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_165" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_165" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/txt">Workload stopped on 2 host(s).</textarea></div><div class="fusion-text fusion-text-164"><p>&nbsp;</p>
<p>This command stops the container and frees the memory. However, the container, Docker image, and model files remain on disk. Therefore, you do not need to re-download to restart — simply run the <code><strong>sparkrun run</strong></code> command from step 3 again.</p>
</div></div></div></div></div><div class="fusion-fullwidth fullwidth-box fusion-builder-row-7 fusion-flex-container nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-6 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:0px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-165"></div></div></div></div></div></p>
<p>The post <a href="https://blog.openzeka.com/en/deepseek-v4-flash-0731-on-2x-dgx-spark/">DeepSeek-V4-Flash-0731 on 2x DGX Spark</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Running GLM-5.2 on Four DGX Sparks</title>
		<link>https://blog.openzeka.com/en/running-glm-5-2-on-four-dgx-sparks/</link>
		
		<dc:creator><![CDATA[Enhar]]></dc:creator>
		<pubDate>Wed, 19 Aug 2026 13:06:31 +0000</pubDate>
				<category><![CDATA[AI Cluster]]></category>
		<guid isPermaLink="false">https://blog-en.openzeka.com/?p=1739</guid>

					<description><![CDATA[<p>In this tutorial, we will run the GLM-5.2 large langua ... Continue Reading→</p>
<p>The post <a href="https://blog.openzeka.com/en/running-glm-5-2-on-four-dgx-sparks/">Running GLM-5.2 on Four DGX Sparks</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-8 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-7 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-166"><p>In this tutorial, we will run the <strong>GLM-5.2</strong> large language model distributed across <strong>4 DGX Spark</strong> units. We will use <strong>sparkrun</strong> for this purpose.</p>
<p>A Spark can be used directly as a computer by connecting a monitor and keyboard, or it can serve as a remote server accessed from another machine. In this tutorial, we will connect to the Spark remotely, install the necessary software, and run the model.</p>
<p>We will use <strong>GLM-5.2</strong> as the language model (a Mixture-of-Experts model with 744 billion parameters, ~40 billion active parameters, int4/int8 quantization), <strong>vLLM</strong> as the inference engine, and <strong>sparkrun</strong> as the management tool. vLLM will load the model’s trained weights into GPU memory and serve an API that accepts requests from outside. sparkrun will manage Docker containers and model distribution across the Sparks via the command line. This automates image synchronization, model transfer, and cluster configuration. Both will run on the <strong>Main Spark</strong> inside Docker containers.</p>
<p><strong>This tutorial consists of three parts:</strong></p>
<ul>
<li><strong>Setup:</strong> Downloading the Docker image and preparing the recipe</li>
<li><strong>Running:</strong> Starting, monitoring, testing, and shutting down the model on 4 Sparks</li>
<li><strong>Benchmark:</strong> Performance measurement at different concurrency levels</li>
</ul>
<p>Throughout this tutorial, the main device is called the <strong>Main Spark</strong>, and the other devices are called <strong>Worker Sparks</strong>.</p>
</div><div class="fusion-text fusion-text-167"><h2><b>Setup</b></h2>
<h3><b>Prerequisites</b></h3>
<p>Previous tutorials covered <b>connecting to the Spark</b>, <b>installing sparkrun</b>, and <b>configuring a multi-Spark cluster</b> step by step. In this tutorial, we assume <b>all these steps have been completed</b> and your setup is <b>ready</b>.</p>
<p>This tutorial requires <b>4 DGX Sparks</b>. We will refer to the main device as <b>&lt;main-spark-ip&gt;</b>, and the other three devices as <b>&lt;worker-spark-ip-1&gt;</b>, <b>&lt;worker-spark-ip-2&gt;</b>, and <b>&lt;worker-spark-ip-3&gt;</b>. Note each device’s <i>IP address</i> beforehand.</p>
<h3><b>1. Downloading the Docker Image</b></h3>
<p><b>GLM-5.2’s DeepSeek Sparse Attention (DSA)</b> architecture requires kernels compiled for <b>GB10 (SM 12.1)</b>. <i>Standard vLLM images</i> do not include these kernels. Therefore, we will use a <b>custom Docker image</b> that contains the necessary patches.</p>
<p>Run the following command to pull the Docker image prepared by <b>OpenZeka</b>:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-166 > .CodeMirror, .fusion-syntax-highlighter-166 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-166 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_166" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_166" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_166" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">docker pull registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest</textarea></div><div class="fusion-text fusion-text-168"><p>&nbsp;</p>
<p>Verify the image:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-167 > .CodeMirror, .fusion-syntax-highlighter-167 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-167 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_167" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_167" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_167" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">docker images registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-168 > .CodeMirror, .fusion-syntax-highlighter-168 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-168 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_168" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_168" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_168" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">REPOSITORY                                            TAG       IMAGE ID        CREATED          SIZE
registry.cordata.ai/spark-cluster/vllm-zatz-dcp       latest    f4c351b85b27    2 minutes ago    29.2GB</textarea></div><div class="fusion-text fusion-text-169"><h3>2. Preparing the Recipe</h3>
<p>The recipe is a YAML file that defines how sparkrun will run the model. The model, Docker image, vLLM flags, and memory settings are consolidated in a single file. For GLM-5.2, this recipe uses tensor parallelism (model weight matrix partitioning) and decode context parallelism (KV cache and attention computation partitioning) across 4 nodes.</p>
<p>Save the recipe file using the following command:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-169 > .CodeMirror, .fusion-syntax-highlighter-169 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-169 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_169" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_169" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_169" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/yaml">cat > ~/glm52-dcp4-cc128.yaml << 'RECIPE'
# GLM-5.2 Int4-Int8Mix TP4 DCP4 128K MTP k=4 cudagraph FULL
# Usage:
#   sparkrun run ~/glm52-dcp4-cc128.yaml              # 4-node cluster (TP=4, DCP=4)
model: QuantTrio/GLM-5.2-Int4-Int8Mix
runtime: vllm-distributed
container: registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest

min_nodes: 4
max_nodes: 4

metadata:
  description: GLM-5.2 Int4-Int8Mix TP4 DCP4 128K MTP k=4 cudagraph FULL
  maintainer: local
  model_params: 744B-MoE-40B-active
  model_dtype: int4-int8mix

defaults:
  port: 8210
  host: 0.0.0.0
  tensor_parallel: 4
  pipeline_parallel: 1
  decode_context_parallel: 4
  gpu_memory_utilization: 0.885
  max_model_len: 131072
  max_num_seqs: 5
  max_num_batched_tokens: 2048
  kv_cache_dtype: fp8_ds_mla
  kv_cache_memory_bytes: 9000000000
  max_cudagraph_capture_size: 32
  load_format: auto
  served_model_name: glm-5.2
  quantization: compressed-tensors
  reasoning_parser: glm45
  tool_call_parser: glm47

env:
  NCCL_IB_TC: "106"
  HF_HUB_OFFLINE: "1"
  VLLM_MARLIN_USE_ATOMIC_ADD: "1"
  TRANSFORMERS_OFFLINE: "1"
  SAFETENSORS_FAST_GPU: "1"
  CUDA_DEVICE_ORDER: "PCI_BUS_ID"
  CUDA_DEVICE_MAX_CONNECTIONS: "32"
  CUTE_DSL_ARCH: "sm_121a"
  TORCH_CUDA_ARCH_LIST: "12.1a"
  VLLM_ALLOW_LONG_MAX_MODEL_LEN: "1"
  VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS: "1800"
  NCCL_MAX_NCHANNELS: "4"
  NCCL_MIN_NCHANNELS: "4"
  NCCL_CROSS_NIC: "1"
  NCCL_CUMEM_ENABLE: "0"
  NCCL_IGNORE_CPU_AFFINITY: "1"
  NCCL_NET_PLUGIN: "none"
  NCCL_IB_MERGE_NICS: "0"
  NCCL_IB_SUBNET_AWARE_ROUTING: "1"
  NCCL_DEBUG: "WARN"
  NCCL_IB_HCA: "rocep1s0f1,roceP2p1s0f1"
  NCCL_SOCKET_IFNAME: "enp1s0f1np1,enP2p1s0f1np1"
  GLOO_SOCKET_IFNAME: "enp1s0f1np1,enP2p1s0f1np1"
  PYTORCH_CUDA_ALLOC_CONF: "expandable_segments:True"
  VLLM_WORKER_MULTIPROC_METHOD: "spawn"
  VLLM_USE_FLASHINFER_SAMPLER: "1"
  VLLM_USE_V2_MODEL_RUNNER: "1"
  VLLM_USE_B12X_SPARSE_INDEXER: "1"
  VLLM_DCP_GLOBAL_TOPK: "1"
  VLLM_DCP_SHARD_DRAFT: "1"
  VLLM_KZ_TRIM_AFTER_LOAD: "1"
  VLLM_USE_B12X_MOE: "0"
  VLLM_USE_B12X_FP8_GEMM: "0"
  VLLM_DISABLE_TP_MQ_BROADCASTER: "1"
  VLLM_ENABLE_PCIE_ALLREDUCE: "0"
  VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS: "0"
  VLLM_SPARSE_INDEXER_MAX_LOGITS_MB: "256"
  USES_B12X: "True"
  FLASHINFER_DISABLE_VERSION_CHECK: "1"
  RAY_memory_usage_threshold: "0.99"
  RAY_memory_monitor_refresh_ms: "0"

command: |
  vllm serve {model} \
      --served-model-name {served_model_name} \
      --trust-remote-code \
      --load-format {load_format} \
      --quantization {quantization} \
      --tensor-parallel-size {tensor_parallel} \
      --pipeline-parallel-size {pipeline_parallel} \
      --decode-context-parallel-size {decode_context_parallel} \
      --dcp-comm-backend ag_rs \
      --dcp-kv-cache-interleave-size 1 \
      --gpu-memory-utilization {gpu_memory_utilization} \
      --max-model-len {max_model_len} \
      --max-num-seqs {max_num_seqs} \
      --max-num-batched-tokens {max_num_batched_tokens} \
      --kv-cache-dtype {kv_cache_dtype} \
      --kv-cache-memory-bytes {kv_cache_memory_bytes} \
      --generation-config vllm \
      --hf-overrides '{"use_index_cache":true,"index_topk_pattern":"FFFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSS"}' \
      --default-chat-template-kwargs '{"clear_thinking":false}' \
      --port {port} \
      --host {host} \
      --compilation-config '{"cudagraph_mode":"FULL","max_cudagraph_capture_size":32}' \
      --no-enable-log-requests \
      --attention-backend B12X_MLA_SPARSE \
      --moe-backend flashinfer_cutlass \
      --reasoning-parser {reasoning_parser} \
      --tool-call-parser {tool_call_parser} \
      --enable-auto-tool-choice \
      --enable-prefix-caching \
      --speculative-config '{"model":"QuantTrio/GLM-5.2-Int4-Int8Mix","method":"mtp","num_speculative_tokens":4,"quantization":"compressed-tensors","moe_backend":"flashinfer_cutlass","draft_attention_backend":"B12X_MLA_SPARSE","draft_sample_method":"probabilistic"}' \
      --long-prefill-token-threshold 2048 \
      --async-scheduling
RECIPE
</textarea></div><div class="fusion-text fusion-text-170"><p>&nbsp;</p>
<p>Verify the file:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-170 > .CodeMirror, .fusion-syntax-highlighter-170 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-170 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_170" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_170" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_170" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">cat ~/glm52-dcp4-cc128.yaml</textarea></div><div class="fusion-text fusion-text-171"><p>&nbsp;</p>
<p>You will see the recipe file you saved.</p>
</div><div class="fusion-text fusion-text-172"><h2>Running</h2>
<h3>1. Clearing the Cache</h3>
<p>Before starting vLLM, clear the filesystem cache on each Spark. The main reason for this is the DGX Spark’s Unified Memory Architecture (UMA): The operating system caches model files read from disk in RAM. vLLM loads the model weights from here into GPU memory. After loading completes, the cached data is not used again, but it is not freed immediately either. On systems with separate memory, this is not important. Since inference runs in GPU memory, RAM utilization does not affect performance. On the Spark, however, since the CPU and GPU share the same RAM, the cache reduces the space available to the GPU. The command clears this cache, providing maximum memory for vLLM.</p>
<p>Clear the cache on the Main Spark:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-171 > .CodeMirror, .fusion-syntax-highlighter-171 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-171 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_171" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_171" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_171" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'

[sudo] password for nvidia:</textarea></div><div class="fusion-text fusion-text-173"><p>&nbsp;</p>
<p>Apply the same cleanup on the Worker Sparks via SSH from the Main Spark:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-172 > .CodeMirror, .fusion-syntax-highlighter-172 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-172 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_172" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_172" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_172" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">for ip in <worker-spark-ip-1> <worker-spark-ip-2> <worker-spark-ip-3>; do
  echo "=== $ip ==="
  ssh nvidia@$ip "sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'"
  echo ""
done

    
=== <worker-spark-ip-1> ===
[sudo] password for nvidia:

=== <worker-spark-ip-2> ===
[sudo] password for nvidia:

=== <worker-spark-ip-3> ===
[sudo] password for nvidia:
</textarea></div><div class="fusion-text fusion-text-174"><h3>2. Pre-launch Checks</h3>
<p>Verify that sparkrun parsed the recipe correctly and that the memory budget is sufficient:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-173 > .CodeMirror, .fusion-syntax-highlighter-173 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-173 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_173" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_173" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_173" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">sparkrun show ~/glm52-dcp4-cc128.yaml</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-174 > .CodeMirror, .fusion-syntax-highlighter-174 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-174 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_174" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_174" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_174" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">Name:         /home/nvidia/glm52-dcp4-cc128.yaml
Description:  GLM-5.2 Int4-Int8Mix TP4 DCP4 128K MTP k=4 cudagraph FULL
Maintainer:   local
Runtime:      vllm-distributed
Model:        QuantTrio/GLM-5.2-Int4-Int8Mix
Container:    registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest
Nodes:        4 - 4

Defaults:
  decode_context_parallel: 4
  gpu_memory_utilization: 0.885
  host: 0.0.0.0
  kv_cache_dtype: fp8_ds_mla
  kv_cache_memory_bytes: 9000000000
  load_format: auto
  max_cudagraph_capture_size: 32
  max_model_len: 131072
  max_num_batched_tokens: 2048
  max_num_seqs: 5
  pipeline_parallel: 1
  port: 8210
  quantization: compressed-tensors
  reasoning_parser: glm45
  served_model_name: glm-5.2
  tensor_parallel: 4
  tool_call_parser: glm47

Environment:
  CUDA_DEVICE_MAX_CONNECTIONS=32
  CUDA_DEVICE_ORDER=PCI_BUS_ID
  CUTE_DSL_ARCH=sm_121a
  FLASHINFER_DISABLE_VERSION_CHECK=1
  GLOO_SOCKET_IFNAME=enp1s0f1np1,enP2p1s0f1np1
  HF_HUB_OFFLINE=1
  NCCL_CROSS_NIC=1
  NCCL_CUMEM_ENABLE=0
  NCCL_DEBUG=WARN
  NCCL_IB_HCA=rocep1s0f1,roceP2p1s0f1
  NCCL_IB_MERGE_NICS=0
  NCCL_IB_SUBNET_AWARE_ROUTING=1
  NCCL_IB_TC=106
  NCCL_IGNORE_CPU_AFFINITY=1
  NCCL_MAX_NCHANNELS=4
  NCCL_MIN_NCHANNELS=4
  NCCL_NET_PLUGIN=none
  NCCL_SOCKET_IFNAME=enp1s0f1np1,enP2p1s0f1np1
  PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
  RAY_memory_monitor_refresh_ms=0
  RAY_memory_usage_threshold=0.99
  SAFETENSORS_FAST_GPU=1
  TORCH_CUDA_ARCH_LIST=12.1a
  TRANSFORMERS_OFFLINE=1
  USES_B12X=True
  VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
  VLLM_DCP_GLOBAL_TOPK=1
  VLLM_DCP_SHARD_DRAFT=1
  VLLM_DISABLE_TP_MQ_BROADCASTER=1
  VLLM_ENABLE_PCIE_ALLREDUCE=0
  VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800
  VLLM_KZ_TRIM_AFTER_LOAD=1
  VLLM_MARLIN_USE_ATOMIC_ADD=1
  VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0
  VLLM_SPARSE_INDEXER_MAX_LOGITS_MB=256
  VLLM_USE_B12X_FP8_GEMM=0
  VLLM_USE_B12X_MOE=0
  VLLM_USE_B12X_SPARSE_INDEXER=1
  VLLM_USE_FLASHINFER_SAMPLER=1
  VLLM_USE_V2_MODEL_RUNNER=1
  VLLM_WORKER_MULTIPROC_METHOD=spawn

Command:
  vllm serve {model} \
    --served-model-name {served_model_name} \
    --trust-remote-code \
    --load-format {load_format} \
    --quantization {quantization} \
    --tensor-parallel-size {tensor_parallel} \
    --pipeline-parallel-size {pipeline_parallel} \
    --decode-context-parallel-size {decode_context_parallel} \
    --dcp-comm-backend ag_rs \
    --dcp-kv-cache-interleave-size 1 \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len {max_model_len} \
    --max-num-seqs {max_num_seqs} \
    --max-num-batched-tokens {max_num_batched_tokens} \
    --kv-cache-dtype {kv_cache_dtype} \
    --kv-cache-memory-bytes {kv_cache_memory_bytes} \
    --generation-config vllm \
    --hf-overrides '{"use_index_cache":true,"index_topk_pattern":"FFFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSS"}' \
    --default-chat-template-kwargs '{"clear_thinking":false}' \
    --port {port} \
    --host {host} \
    --compilation-config '{"cudagraph_mode":"FULL","max_cudagraph_capture_size":32}' \
    --no-enable-log-requests \
    --attention-backend B12X_MLA_SPARSE \
    --moe-backend flashinfer_cutlass \
    --reasoning-parser {reasoning_parser} \
    --tool-call-parser {tool_call_parser} \
    --enable-auto-tool-choice \
    --enable-prefix-caching \
    --speculative-config '{"model":"QuantTrio/GLM-5.2-Int4-Int8Mix","method":"mtp","num_speculative_tokens":4,"quantization":"compressed-tensors","moe_backend":"flashinfer_cutlass","draft_attention_backend":"B12X_MLA_SPARSE","draft_sample_method":"probabilistic"}' \
    --long-prefill-token-threshold 2048 \
    --async-scheduling
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.

VRAM Estimation:
  Model dtype:      int4_int8mix
  Model params:     202,740,507,008
  KV cache dtype:   fp8_ds_mla
  Architecture:     78 layers, 64 KV heads, 192 head_dim
  Model weights:    0.00 GB
  Tensor parallel:  4
  Per-GPU total:    0.00 GB
  DGX Spark fit:    YES

  GPU Memory Budget:
    gpu_memory_utilization: 88%
    Usable GPU memory:     107.1 GB (121 GB x 88%)
    Available for KV:      107.1 GB
  Warning: Unknown dtype 'int4_int8mix'; cannot estimate model weight VRAM
  Warning: Unknown KV cache dtype 'fp8_ds_mla'
</textarea></div><div class="fusion-text fusion-text-175"><p>&nbsp;</p>
<p>sparkrun parsed the recipe correctly, selected vllm-distributed, and confirmed with ‘DGX Spark fit: YES’. You can ignore the unknown dtype warnings — these are b12x/DCP-specific types that sparkrun does not recognize.</p>
<p>Now preview the launch plan:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-175 > .CodeMirror, .fusion-syntax-highlighter-175 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-175 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_175" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_175" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_175" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">sparkrun run ~/glm52-dcp4-cc128.yaml --dry-run</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-176 > .CodeMirror, .fusion-syntax-highlighter-176 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-176 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_176" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_176" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_176" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">Warning: metadata.model_params '744B-MoE-40B-active' is not a valid parameter count
Warning: metadata.model_dtype 'int4-int8mix' is not a recognized dtype
sparkrun v0.2.40

Runtime:   vllm-distributed
Image:     registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest
Model:     QuantTrio/GLM-5.2-Int4-Int8Mix
Mode:      cluster (4 nodes)
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.

VRAM Estimation:
  Model dtype:      int4_int8mix
  Model params:     202,740,507,008
  KV cache dtype:   fp8_ds_mla
  Architecture:     78 layers, 64 KV heads, 192 head_dim
  Model weights:    0.00 GB
  Tensor parallel:  4
  Per-GPU total:    0.00 GB
  DGX Spark fit:    YES

  GPU Memory Budget:
    gpu_memory_utilization: 88%
    Usable GPU memory:     107.1 GB (121 GB x 88%)
    Available for KV:      107.1 GB
  Warning: Unknown dtype 'int4_int8mix'; cannot estimate model weight VRAM
  Warning: Unknown KV cache dtype 'fp8_ds_mla'

Hosts:     default cluster 'default'
  Head:    <main-spark-ip>
  Workers: <worker-spark-ip-1>, <worker-spark-ip-2>, <worker-spark-ip-3>

[1/6] Preparing
  done (0.0s)
[2/6] Building image — skipped (no builder)
[3/6] Distributing resources
  Distributing image registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest to 4 host(s)
  Distributing model QuantTrio/GLM-5.2-Int4-Int8Mix to 4 host(s)
  Model synced to 4 host(s)
  done (0.1s)
[4/6] Syncing tuning configs
  done (0.0s)
[5/6] Launching vllm runtime
  Step 1/7: Cleaning up existing containers
  Step 2/7: Detecting InfiniBand
  Step 3/7: Detecting head node IP
  Step 4/7: Launching containers
  Step 5/7: Running pre-serve hooks
  Step 6/7: Starting head node serve
  Step 7/7: Starting worker nodes
  done (0.1s)
Cluster:   sparkrun_1e20fe4f386b

Serve command:
  vllm serve QuantTrio/GLM-5.2-Int4-Int8Mix \
      --served-model-name glm-5.2 \
      --trust-remote-code \
      --load-format auto \
      --quantization compressed-tensors \
      --tensor-parallel-size 4 \
      --pipeline-parallel-size 1 \
      --decode-context-parallel-size 4 \
      --dcp-comm-backend ag_rs \
      --dcp-kv-cache-interleave-size 1 \
      --gpu-memory-utilization 0.885 \
      --max-model-len 131072 \
      --max-num-seqs 5 \
      --max-num-batched-tokens 2048 \
      --kv-cache-dtype fp8_ds_mla \
      --kv-cache-memory-bytes 9000000000 \
      --generation-config vllm \
      --hf-overrides '{"use_index_cache":true,"index_topk_pattern":"FFFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSS"}' \
      --default-chat-template-kwargs '{"clear_thinking":false}' \
      --port 8210 \
      --host 0.0.0.0 \
      --compilation-config '{"cudagraph_mode":"FULL","max_cudagraph_capture_size":32}' \
      --no-enable-log-requests \
      --attention-backend B12X_MLA_SPARSE \
      --moe-backend flashinfer_cutlass \
      --reasoning-parser glm45 \
      --tool-call-parser glm47 \
      --enable-auto-tool-choice \
      --enable-prefix-caching \
      --speculative-config '{"model":"QuantTrio/GLM-5.2-Int4-Int8Mix","method":"mtp","num_speculative_tokens":4,"quantization":"compressed-tensors","moe_backend":"flashinfer_cutlass","draft_attention_backend":"B12X_MLA_SPARSE","draft_sample_method":"probabilistic"}' \
      --long-prefill-token-threshold 2048 \
      --async-scheduling

[6/6] Post-launch hooks — skipped
</textarea></div><div class="fusion-text fusion-text-176"><p>&nbsp;</p>
<p>The dry run succeeded. The recipe is valid, ‘Mode: cluster (4 nodes)’ is shown, and the serve command includes all flags. If the model is not installed on the Spark, sparkrun will automatically download it from Hugging Face on first launch. The dry run does not trigger this download; downloading only happens during an actual launch.</p>
</div><div class="fusion-text fusion-text-177"><h3>3. Starting the Model</h3>
<p>Now let’s start the model:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-177 > .CodeMirror, .fusion-syntax-highlighter-177 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-177 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_177" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_177" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_177" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">sparkrun run ~/glm52-dcp4-cc128.yaml --no-follow</textarea></div><div class="fusion-text fusion-text-178"><p>&nbsp;</p>
<p>The<em><strong> &#8211;no-follow</strong></em> flag makes sparkrun return to the command line after launching the containers. sparkrun automatically synchronizes the image to the Workers (skips if same ID), downloads the model to the head node and distributes it to the Workers (skips if already present), configures NCCL for the CX-7 interfaces, and launches a container on each Spark.</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-178 > .CodeMirror, .fusion-syntax-highlighter-178 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-178 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_178" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_178" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_178" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">Warning: metadata.model_params '744B-MoE-40B-active' is not a valid parameter count
Warning: metadata.model_dtype 'int4-int8mix' is not a recognized dtype
sparkrun v0.2.40

Runtime:   vllm-distributed
Image:     registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest
Model:     QuantTrio/GLM-5.2-Int4-Int8Mix
Mode:      cluster (4 nodes)
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.

VRAM Estimation:
  Model dtype:      int4_int8mix
  Model params:     202,740,507,008
  KV cache dtype:   fp8_ds_mla
  Architecture:     78 layers, 64 KV heads, 192 head_dim
  Model weights:    0.00 GB
  Tensor parallel:  4
  Per-GPU total:    0.00 GB
  DGX Spark fit:    YES

  GPU Memory Budget:
    gpu_memory_utilization: 88%
    Usable GPU memory:     107.1 GB (121 GB x 88%)
    Available for KV:      107.1 GB
  Warning: Unknown dtype 'int4_int8mix'; cannot estimate model weight VRAM
  Warning: Unknown KV cache dtype 'fp8_ds_mla'

Hosts:     default cluster 'default'
  Head:    <main-spark-ip>
  Workers: <worker-spark-ip-1>, <worker-spark-ip-2>, <worker-spark-ip-3>

[1/6] Preparing
  done (0.1s)
[2/6] Building image — skipped (no builder)
[3/6] Distributing resources
  Distributing image registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest to 4 host(s)
  Container image stale on 3 of 4 host(s), syncing
  Distributing model QuantTrio/GLM-5.2-Int4-Int8Mix to 4 host(s)
Fetching 142 files:   0%|          | 0/142 [00:00<?, ?it/s]
Downloading (.…):   0%|          | 0.00/378G [00:00<?, ?iB/s]
Downloading (.…): 100%|██████████| 378G/378G [42:18<00:00, 149MiB/s]
Fetching 142 files: 100%|██████████| 142/142 [42:19<00:00, 17.84it/s]
  Model synced to 4 host(s)
  done (2418.5s)
[4/6] Syncing tuning configs
  done (0.0s)
[5/6] Launching vllm runtime
  Step 1/7: Cleaning up existing containers
  Step 2/7: Detecting InfiniBand
  Step 3/7: Detecting head node IP
  Step 4/7: Launching containers
  Step 5/7: Running pre-serve hooks
  Step 6/7: Starting head node serve
  Step 7/7: Starting worker nodes
  done (52.1s)
Cluster:   sparkrun_1e20fe4f386b

Serve command:
  vllm serve QuantTrio/GLM-5.2-Int4-Int8Mix \
      --served-model-name glm-5.2 \
      --trust-remote-code \
      --load-format auto \
      --quantization compressed-tensors \
      --tensor-parallel-size 4 \
      --pipeline-parallel-size 1 \
      --decode-context-parallel-size 4 \
      --dcp-comm-backend ag_rs \
      --dcp-kv-cache-interleave-size 1 \
      --gpu-memory-utilization 0.885 \
      --max-model-len 131072 \
      --max-num-seqs 5 \
      --max-num-batched-tokens 2048 \
      --kv-cache-dtype fp8_ds_mla \
      --kv-cache-memory-bytes 9000000000 \
      --generation-config vllm \
      --hf-overrides '{"use_index_cache":true,"index_topk_pattern":"FFFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSS"}' \
      --default-chat-template-kwargs '{"clear_thinking":false}' \
      --port 8210 \
      --host 0.0.0.0 \
      --compilation-config '{"cudagraph_mode":"FULL","max_cudagraph_capture_size":32}' \
      --no-enable-log-requests \
      --attention-backend B12X_MLA_SPARSE \
      --moe-backend flashinfer_cutlass \
      --reasoning-parser glm45 \
      --tool-call-parser glm47 \
      --enable-auto-tool-choice \
      --enable-prefix-caching \
      --speculative-config '{"model":"QuantTrio/GLM-5.2-Int4-Int8Mix","method":"mtp","num_speculative_tokens":4,"quantization":"compressed-tensors","moe_backend":"flashinfer_cutlass","draft_attention_backend":"B12X_MLA_SPARSE","draft_sample_method":"probabilistic"}' \
      --long-prefill-token-threshold 2048 \
      --async-scheduling

Runtime versions:
  cuda:      13.2
  nccl:      (2, 28, 9)
  python:    3.12.3
  torch:     2.11.0+cu130
  vllm:      0.1.dev17863+ge232d2623.d20260713

[6/6] Post-launch hooks — skipped
</textarea></div><div class="fusion-text fusion-text-179"><p>&nbsp;</p>
<p>sparkrun completed all 6 steps successfully. <strong>‘Mode: cluster (4 nodes)’</strong> is shown. sparkrun synchronized the image to 3 Workers over the CX-7 network and downloaded the model from Hugging Face and distributed it to 4 nodes. All flags in the serve command resolved correctly.</p>
</div><div class="fusion-text fusion-text-180"><h3>4. Monitoring the Startup Process</h3>
<p>After model downloading completes, it may take a few minutes for vLLM to become ready for serving. During this process, vLLM loads the model weights into GPU memory, compiles GPU kernels, and allocates memory for inference.</p>
<p>To view vLLM logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-179 > .CodeMirror, .fusion-syntax-highlighter-179 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-179 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_179" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_179" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_179" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">sparkrun logs ~/glm52-dcp4-cc128.yaml</textarea></div><div class="fusion-text fusion-text-181"><p>&nbsp;</p>
<p>You will see the following in the logs:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-180 > .CodeMirror, .fusion-syntax-highlighter-180 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-180 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_180" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_180" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_180" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">(APIServer pid=43) INFO 07-31 06:56:57 [api_utils.py:339]  ▄▄ ▄█ █     █     █ ▀▄▀ █  version 0.1.dev17863+ge232d2623.d20260713
(APIServer pid=43) INFO 07-31 06:56:57 [api_utils.py:339]   █▄█▀ █     █     █     █  model   QuantTrio/GLM-5.2-Int4-Int8Mix

(APIServer pid=43) INFO 07-31 06:57:05 [model.py:598] Resolved architecture: GlmMoeDsaForCausalLM
(APIServer pid=43) INFO 07-31 06:57:05 [model.py:1725] Using max model len 131072
(APIServer pid=43) INFO 07-31 06:57:05 [cache.py:279] Using fp8_ds_mla data type to store kv cache.

(APIServer pid=43) INFO 07-31 06:57:12 [model.py:598] Resolved architecture: DeepSeekMTPModel
(APIServer pid=43) INFO 07-31 06:57:12 [speculative.py:1072] Overriding draft model max model len from 1048576 to 131072
(APIServer pid=43) INFO 07-31 06:57:12 [scheduler.py:252] Chunked prefill is enabled with max_num_batched_tokens=2048.
(APIServer pid=43) INFO 07-31 06:57:12 [vllm.py:1051] Asynchronous scheduling is enabled.

(EngineCore pid=113) INFO 07-31 06:57:20 [core.py:114] Initializing a V1 LLM engine (v0.1.dev17863+ge232d2623.d20260713) with config: ...
(EngineCore pid=113) INFO 07-31 06:57:20 [multiproc_executor.py:140] DP group leader: node_rank=0, master_addr=<main-spark-ip>, world_size=4

(Worker pid=129) INFO 07-31 06:57:28 [parallel_state.py:1568] world_size=4 rank=0 distributed_init_method=tcp://<main-spark-ip>:25000 backend=nccl

(Worker_TP0_DCP0 pid=129) INFO 07-31 06:58:07 [model_runner.py:312] Loading model from scratch...
(Worker_TP0_DCP0 pid=129) INFO 07-31 06:58:07 [cuda.py:404] Using AttentionBackendEnum.B12X_MLA_SPARSE backend.
(Worker_TP0_DCP0 pid=129) INFO 07-31 06:58:10 [selector.py:166] Using FLASH_ATTN MLA prefill backend.
(Worker_TP0_DCP0 pid=129) INFO 07-31 06:58:10 [compressed_tensors_moe.py:155] Using CompressedTensorsWNA16MarlinMoEMethod

(Worker_TP0_DCP0 pid=129) INFO 07-31 06:58:14 [weight_utils.py:914] Filesystem type for checkpoints: EXT4. Checkpoint size: 377.63 GiB. Available RAM: 16.52 GiB.
(Worker_TP0_DCP0 pid=129) Loading safetensors checkpoint shards: 100% Completed | 128/128 [05:14<00:00,  2.46s/it]
(Worker_TP0_DCP0 pid=129) INFO 07-31 07:03:29 [default_loader.py:451] Loading weights took 314.57 seconds

(Worker_TP0_DCP0 pid=129) INFO 07-31 07:03:40 [weight_utils.py:914] Filesystem type for checkpoints: EXT4. Checkpoint size: 9.36 GiB. Available RAM: 15.73 GiB.
(Worker_TP0_DCP0 pid=129) Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:07<00:00,  1.90s/it]
(Worker_TP0_DCP0 pid=129) INFO 07-31 07:03:48 [default_loader.py:451] Loading weights took 7.63 seconds
(Worker_TP0_DCP0 pid=129) INFO 07-31 07:03:56 [model_runner.py:333] Model loading took 96.71 GiB and 349.827079 seconds

(Worker_TP0_DCP0 pid=129) INFO 07-31 07:04:10 [backends.py:1089] Using cache directory: /tmp/.cache/vllm/torch_compile_cache/.../backbone for vLLM's torch.compile
(Worker_TP0_DCP0 pid=129) INFO 07-31 07:04:36 [monitor.py:53] torch.compile took 38.46 s in total
(Worker_TP0_DCP0 pid=129) INFO 07-31 07:05:15 [backends.py:1089] Using cache directory: /tmp/.cache/vllm/torch_compile_cache/.../eagle_head for vLLM's torch.compile
(Worker_TP0_DCP0 pid=129) INFO 07-31 07:05:20 [monitor.py:53] torch.compile took 6.47 s in total

(Worker_TP0_DCP0 pid=129) INFO 07-31 07:05:25 [gpu_worker.py:407] Initial free memory 113.45 GiB, reserved 8.38 GiB memory for KV Cache as specified by kv_cache_memory_bytes config and skipped memory profiling.

(EngineCore pid=113) INFO 07-31 07:05:25 [kv_cache_utils.py:2152] GPU KV cache size: 657,664 tokens
(EngineCore pid=113) INFO 07-31 07:05:25 [kv_cache_utils.py:2153] Maximum concurrency for 131,072 tokens per request: 5.02x

(Worker_TP0_DCP0 pid=129) WARNING 07-31 07:05:25 [compilation.py:1361] CUDAGraphMode.FULL is not supported with B12xNonCompressedIndexerBackend backend; setting cudagraph_mode=FULL_AND_PIECEWISE
(Worker_TP0_DCP0 pid=129) Capturing CUDA graphs (PIECEWISE): 100%|██████████| 6/6 [00:01<00:00,  4.63it/s]
(Worker_TP0_DCP0 pid=129) Capturing CUDA graphs (FULL): 100%|██████████| 5/5 [00:06<00:00,  1.39it/s]
(Worker_TP0_DCP0 pid=129) INFO 07-31 07:05:39 [model_runner.py:766] Graph capturing finished in 14 secs, took 0.27 GiB

(APIServer pid=43) INFO 07-31 07:06:04 [parser_manager.py:37] "auto" tool choice has been enabled.
(APIServer pid=43) INFO 07-31 07:06:07 [api_server.py:576] Starting vLLM server on http://0.0.0.0:8210
(APIServer pid=43) INFO:     Started server process [43]
(APIServer pid=43) INFO:     Waiting for application startup.
(APIServer pid=43) INFO:     Application startup complete.
</textarea></div><div class="fusion-text fusion-text-182"><p>&nbsp;</p>
<p>If you see the <b>Application startup complete.</b> line, the <b>model server is ready</b>. Press <b>Ctrl+C</b> to stop watching logs. The server will <b>continue running in the background</b>.</p>
<p>The startup process takes <b>approximately 10 minutes</b>. During this process:</p>
<ul>
<li>The architecture is resolved as <b>GlmMoeDsaForCausalLM</b> (<i>DeepSeek Sparse Attention + MoE</i>).</li>
<li>The <b>MTP draft model (DeepSeekMTPModel)</b> is initialized — <b>speculative decoding</b> is active (<b>k=4</b>).</li>
<li>The <b>B12X_MLA_SPARSE</b> attention backend is selected — <i>DeepGEMM</i> is skipped.</li>
<li><b>128 safetensors shards</b> are loaded (<b>~5 minutes</b>, <b>96.71 GiB per node</b>).</li>
<li>The <b>draft model</b> is loaded as <b>4 shards</b> (<b>~8 seconds</b>).</li>
<li><b>torch.compile</b> takes <b>~38 seconds</b> (backbone) + <b>~6 seconds</b> (eagle head).</li>
<li><b>CUDA graph capture</b> succeeds in <b>FULL</b> + <b>PIECEWISE</b> modes (<b>14 seconds</b>).</li>
<li><b>KV cache:</b> <b>8.38 GiB</b>, <b>657,664 tokens</b>, <b>5.02× concurrency</b> for <b>131K tokens</b>.</li>
</ul>
</div><div class="fusion-text fusion-text-183"><h3>5. Testing the Model</h3>
<p>First, verify that the server is running:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-181 > .CodeMirror, .fusion-syntax-highlighter-181 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-181 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_181" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_181" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_181" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">curl -s -o /dev/null -w "HTTP %{http_code}" http://<main-spark-ip>:8210/health</textarea></div><div class="fusion-text fusion-text-184"><p>&nbsp;</p>
<p>An ‘HTTP 200’ response indicates the server is healthy. Now check the sparkrun container status:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-182 > .CodeMirror, .fusion-syntax-highlighter-182 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-182 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_182" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_182" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_182" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">sparkrun status</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-183 > .CodeMirror, .fusion-syntax-highlighter-183 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-183 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_183" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_183" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_183" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">Job: /home/nvidia/glm52-dcp4-cc128.yaml  (tp=4, pp=1)  [1e20fe4f386b]  (4 container(s))
  node_0     <main-spark-ip>                             Up 35 minutes             registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest
  node_1     <worker-spark-ip-1>                         Up 35 minutes             registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest
  node_2     <worker-spark-ip-2>                         Up 35 minutes             registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest
  node_3     <worker-spark-ip-3>                         Up 35 minutes             registry.cordata.ai/spark-cluster/vllm-zatz-dcp:latest
  logs: sparkrun logs 1e20fe4f386b
  stop: sparkrun stop 1e20fe4f386b

Total: 4 container(s) across 4 host(s)</textarea></div><div class="fusion-text fusion-text-185"><p>&nbsp;</p>
<p>All four containers are running.</p>
<p>List the registered models:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-184 > .CodeMirror, .fusion-syntax-highlighter-184 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-184 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_184" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_184" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_184" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">curl -s http://localhost:8210/v1/models | python3 -m json.tool</textarea></div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-185 > .CodeMirror, .fusion-syntax-highlighter-185 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-185 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_185" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_185" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_185" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">{
    "object": "list",
    "data": [
        {
            "id": "glm-5.2",
            "object": "model",
            "created": 1784789897,
            "owned_by": "vllm",
            "root": "QuantTrio/GLM-5.2-Int4-Int8Mix",
            "parent": null,
            "max_model_len": 131072
        }
    ]
}</textarea></div><div class="fusion-text fusion-text-186"><p>&nbsp;</p>
<p>The model is registered as <em><strong>glm-5.2.</strong></em></p>
<p>Now test the model with a simple “What is 2+2?” prompt:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-186 > .CodeMirror, .fusion-syntax-highlighter-186 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-186 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_186" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_186" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_186" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">curl -s http://localhost:8210/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.2",
    "messages": [{"role": "user", "content": "2+2 kac eder? Tek cumle?"}],
    "max_tokens": 200
  }' | python3 -m json.tool
</textarea></div><div class="fusion-text fusion-text-187"><p>&nbsp;</p>
<p>The simplified response will be as follows. In the <strong><em>content</em></strong> field (the model’s response), you will see “2+2, 4 eder.” From this, you can see that the model performed the addition correctly. Additionally, the response has a <strong><em>reasoning</em></strong> field. This field contains the model’s thinking process:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-187 > .CodeMirror, .fusion-syntax-highlighter-187 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-187 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_187" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_187" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_187" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="hopscotch" data-mode="text/x-sh">{
    "id": "chatcmpl-aefb2db7bdfbc5c9",
    "object": "chat.completion",
    "created": 1785483144,
    "model": "glm-5.2",
    "choices": [
        {
            "index": 0,
            "message": {
                "role": "assistant",
                "content": "2+2 dört eder.",
                "refusal": null,
                "annotations": null,
                "audio": null,
                "function_call": null,
                "tool_calls": [],
                "reasoning": "1.  **Analyze the Request:**\n    *   Question: \"2+2 kac eder?\" (What is 2+2?)\n    *   Constraint: \"Tek cumle.\" (Single sentence.)\n    *   Language: Turkish.\n\n2.  **Formulate the Answer:**\n    *   2 + 2 = 4.\n    *   Turkish translation: \"2 artı 2 dört eder.\" or simply \"2+2 dört eder.\" or \"İki artı iki dört eder.\"\n\n3.  **Check Constraints:**\n    *   Is it a single sentence? Yes.\n\n4.  **Select the Best Option:**\n    *   \"2+2 dört eder.\" is direct, accurate, and perfectly fits the single sentence constraint."
            },
            "logprobs": null,
            "finish_reason": "stop",
            "stop_reason": 154827,
            "token_ids": null,
            "routed_experts": null
        }
    ],
    "service_tier": null,
    "system_fingerprint": "vllm-0.1.dev17863+ge232d2623.d20260713-tp4-d57ceda6",
    "usage": {
        "prompt_tokens": 25,
        "total_tokens": 198,
        "completion_tokens": 173,
        "prompt_tokens_details": null
    },
    "prompt_logprobs": null,
    "prompt_token_ids": null,
    "prompt_text": null,
    "kv_transfer_params": null
}
</textarea></div><div class="fusion-text fusion-text-188"><p>&nbsp;</p>
<p>The LLM is now running and serving from the Main Spark’s port <strong>8210</strong>. We can ask GLM-5.2 questions and receive responses. You can use the model with a web interface (e.g., Open WebUI) or an agent architecture (e.g., OpenCode).</p>
</div><div class="fusion-text fusion-text-189"><h3>6. Shutdown</h3>
<p>When you are done, stop the model:</p>
</div><style type="text/css" scopped="scopped">.fusion-syntax-highlighter-188 > .CodeMirror, .fusion-syntax-highlighter-188 > .CodeMirror .CodeMirror-gutters {background-color:#000000;}</style><div class="fusion-syntax-highlighter-container fusion-syntax-highlighter-188 fusion-syntax-highlighter-theme-dark" style="opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);"><div class="syntax-highlighter-copy-code"><span class="syntax-highlighter-copy-code-title" data-id="fusion_syntax_highlighter_188" style="font-size:14px;">Copy to Clipboard</span></div><label for="fusion_syntax_highlighter_188" class="screen-reader-text">Syntax Highlighter</label><textarea class="fusion-syntax-highlighter-textarea" id="fusion_syntax_highlighter_188" data-readOnly="nocursor" data-lineNumbers="" data-lineWrapping="" data-theme="oceanic-next" data-mode="text/x-sh">sparkrun stop ~/glm52-dcp4-cc128.yaml

Workload stopped on 4 host(s).</textarea></div><div class="fusion-text fusion-text-190"><p>&nbsp;</p>
<p>This command stops the container and frees the memory space. However, the container, Docker image, and model files remain on disk. This way, you do not need to download them again to restart. Simply run the <strong>sparkrun run</strong> command from step 3 again.</p>
</div><div class="fusion-text fusion-text-191"><h2><b>Benchmark</b></h2>
<p>We tested the <b>GLM-5.2</b> model at different <b>concurrency levels</b>. The measurements recorded average <b>TTFT (Time to First Token)</b> — <i>the time until the first token is produced</i> — and <b>TPS (Tokens Per Second)</b> — <i>the number of tokens generated per second</i>. As can be seen, the value exceeded <b>27 tokens per second</b>.</p>
<table>
<thead>
<tr>
<th><b>Concurrency</b></th>
<th><b>Avg TTFT (ms)</b></th>
<th><b>Avg TPS (tok/s)</b></th>
</tr>
</thead>
<tbody>
<tr>
<td><b>1</b></td>
<td>490</td>
<td><b>27.4</b></td>
</tr>
<tr>
<td><b>2</b></td>
<td>717</td>
<td>20.0</td>
</tr>
<tr>
<td><b>4</b></td>
<td>955</td>
<td>14.3</td>
</tr>
<tr>
<td><b>8</b></td>
<td>6560</td>
<td>8.5</td>
</tr>
</tbody>
</table>
<p>&nbsp;</p>
<p>The measurements were taken using the <b>CordatusAI LLM Benchmark Tool</b>. This is a <i>benchmarking application</i> developed by <b>CordatusAI</b> that tests <b>LLM servers</b> with <b>OpenAI-compatible APIs</b>. Below, you can see the <b>web interface</b> of our benchmark tool and the various <b>charts</b> it produced:</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-38 hover-type-none"><img decoding="async" width="1024" height="579" title="glm_benchmark_3" src="https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-1024x579.webp" alt class="img-responsive wp-image-1744" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-200x113.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-300x170.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-400x226.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-600x339.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-768x434.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-800x452.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-1024x579.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-1200x678.webp 1200w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3-1536x868.webp 1536w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_3.webp 1850w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-39 hover-type-none"><img decoding="async" width="700" height="500" title="glm_benchmark_1" src="https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_1.webp" alt class="img-responsive wp-image-1742" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_1-200x143.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_1-300x214.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_1-400x286.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_1-600x429.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_1.webp 700w" sizes="(max-width: 640px) 100vw, 700px" /></span></div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-40 hover-type-none"><img decoding="async" width="700" height="500" title="glm_benchmark_2" src="https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_2.webp" alt class="img-responsive wp-image-1743" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_2-200x143.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_2-300x214.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_2-400x286.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_2-600x429.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/glm_benchmark_2.webp 700w" sizes="(max-width: 640px) 100vw, 700px" /></span></div></div></div></div></div><div class="fusion-fullwidth fullwidth-box fusion-builder-row-9 fusion-flex-container nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-8 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-blend:overlay;--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:0px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-flex-justify-content-flex-start fusion-content-layout-column"></div></div></div></div></p>
<p>The post <a href="https://blog.openzeka.com/en/running-glm-5-2-on-four-dgx-sparks/">Running GLM-5.2 on Four DGX Sparks</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Building a High-Speed AI Cluster with 4 DGX Spark Systems</title>
		<link>https://blog.openzeka.com/en/building-a-high-speed-ai-cluster-with-4-dgx-spark-systems/</link>
		
		<dc:creator><![CDATA[Enhar]]></dc:creator>
		<pubDate>Thu, 30 Jul 2026 08:10:36 +0000</pubDate>
				<category><![CDATA[AI Cluster]]></category>
		<category><![CDATA[Generative AI]]></category>
		<guid isPermaLink="false">https://blog-en.openzeka.com/?p=1715</guid>

					<description><![CDATA[<p>For large AI models, the limiting factor is often not  ... Continue Reading→</p>
<p>The post <a href="https://blog.openzeka.com/en/building-a-high-speed-ai-cluster-with-4-dgx-spark-systems/">Building a High-Speed AI Cluster with 4 DGX Spark Systems</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-10 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1331.2px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-9 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-192"><p>For large AI models, the limiting factor is often not just compute performance, but memory capacity. When a model no longer fits within the memory of a single system, the workload must be distributed across multiple nodes. However, multi-node deployments are not used solely to accommodate models that are too large for one system. Even when a model can run on a single DGX Spark, using two or four systems with the right parallelism strategy and runtime configuration can increase aggregate token throughput and concurrent request-processing capacity.</p>
<p>There is an important distinction here: adding more nodes does not automatically improve performance in every scenario. The model architecture, inference engine, parallelism strategy, request pattern, and inter-node communication overhead all have a direct impact on the outcome. The real benefit of a multi-node system therefore comes not from the hardware count alone, but from optimizing the compute and communication layers together.</p>
<p>In this project, four NVIDIA DGX Spark systems were connected over a high-speed network and configured as a single AI cluster. The goal was to run models that could not be hosted on one Spark alone, while also increasing token throughput for suitable workloads.</p>
<p>The 10GbE RJ45 management network was separated from the 200GbE compute network provided by the ConnectX-7 interfaces. When both ConnectX-7 interfaces on the compute fabric were used together, approximately <strong>200 Gbps of aggregate RDMA bandwidth</strong> was achieved between the Spark systems, with latency measured at<strong> 1–3 microseconds.</strong> The infrastructure was then validated under a real workload by running GLM-5.2 across all four nodes.</p>
</div><div class="fusion-text fusion-text-193"><h3>Why Build a Cluster?</h3>
<p>When a model does not fit within the memory of a single GPU or system, several parallelism strategies can be used.</p>
<p>Tensor parallelism splits large tensors and computations within an individual model layer across multiple GPUs. Each GPU processes part of the same layer, and the intermediate results are exchanged between devices. This approach is particularly useful when a single layer or the model weights cannot fit into the memory of one GPU.</p>
<p>Pipeline parallelism divides the model into groups of layers or sequential stages. The first layers run on one GPU or node, while later layers run on another. Data, typically organized into micro-batches, moves through these stages like items along a production pipeline. Pipeline parallelism is especially useful in multi-node deployments because different parts of the model can be placed on separate devices.</p>
<p>Tensor and pipeline parallelism can also be combined. The most appropriate strategy depends on the model architecture, memory requirements, and the capabilities of the inference stack.</p>
<p>Because GPUs continuously exchange data during distributed execution, compute performance alone is not enough. If the network is slow, GPUs spend time waiting for data from other nodes instead of performing calculations. Cluster performance must therefore be evaluated as a combination of GPU compute capability and inter-node communication performance.</p>
</div><div class="fusion-text fusion-text-194"><h3>Architecture</h3>
<p>The cluster was built on two physically separate networks:</p>
<ul>
<li><strong>Management network (10GbE):</strong> SSH access, system administration, software updates, and NAS traffic</li>
<li><strong>Compute network (200GbE):</strong> RDMA and distributed model communication</li>
</ul>
<p>This separation prevents intensive compute traffic from interfering with routine management and storage operations. The main components and their roles are summarized below:</p>
<div class="table-1">
<table width="100%">
<thead>
<tr>
<th align="left">Component</th>
<th align="left">Role</th>
</tr>
</thead>
<tbody>
<tr>
<td>4× NVIDIA DGX Spark</td>
<td>Compute nodes running the distributed model workload</td>
</tr>
<tr>
<td>MikroTik CRS312</td>
<td>10GbE management network and NAS connectivity</td>
</tr>
<tr>
<td>MikroTik CRS812</td>
<td>200GbE compute network</td>
</tr>
<tr>
<td>ASUSTOR AS6808T</td>
<td>Shared storage for models and datasets</td>
</tr>
<tr>
<td>sparkrun</td>
<td>Cluster setup and multi-node workload management</td>
</tr>
</tbody>
</table>
</div>
<p>&nbsp;</p>
<p>The 400G QSFP-DD ports on the CRS812 were split into two independent 200G QSFP56 links using passive breakout cables. This allowed a single physical switch port to provide separate 200G connections to two Spark systems.</p>
</div><div class="fusion-image-element " style="text-align:center;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-41 hover-type-none"><img decoding="async" width="634" height="92" title="06-breakout-cable" src="https://blog.openzeka.com/en/wp-content/uploads/2026/07/06-breakout-cable.webp" alt class="img-responsive wp-image-1717" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/07/06-breakout-cable-200x29.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/06-breakout-cable-300x44.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/06-breakout-cable-400x58.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/06-breakout-cable-600x87.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/06-breakout-cable.webp 634w" sizes="(max-width: 640px) 100vw, 634px" /></span></div><div class="fusion-text fusion-text-195"><p>&nbsp;</p>
<p>&nbsp;</p>
</div><div class="fusion-image-element " style="text-align:center;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-42 hover-type-none"><img decoding="async" width="638" height="227" title="07-connectx7-ports" src="https://blog.openzeka.com/en/wp-content/uploads/2026/07/07-connectx7-ports.webp" alt class="img-responsive wp-image-1718" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/07/07-connectx7-ports-200x71.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/07-connectx7-ports-300x107.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/07-connectx7-ports-400x142.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/07-connectx7-ports-600x213.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/07-connectx7-ports.webp 638w" sizes="(max-width: 640px) 100vw, 638px" /></span></div><div class="fusion-text fusion-text-196"><h3></h3>
<h3>Why RDMA and RoCEv2?</h3>
<p>With conventional TCP/IP communication, data passes through the operating system networking stack. This process can introduce CPU overhead, memory copies, and additional latency. In large-model workloads, where high volumes of data are continuously exchanged between nodes, that overhead becomes more significant.</p>
<p>RDMA, or Remote Direct Memory Access, enables data to be transferred from the memory of one system to another while reducing the load on the CPU and operating system networking stack. This can deliver lower latency and higher effective bandwidth.</p>
<p>This cluster uses RoCEv2, or RDMA over Converged Ethernet version 2, to carry RDMA traffic over an Ethernet fabric. Ethernet is inherently a lossy network, however, so packet loss and congestion behavior must be carefully managed for RoCEv2 to operate efficiently.</p>
<p>The following settings were therefore applied to the compute network:</p>
<ul>
<li><strong>MTU 9000:</strong> Reduces packet-processing overhead by carrying large transfers in fewer frames</li>
<li><strong>DSCP 26 → TC3:</strong> Places RoCE traffic into a dedicated traffic class</li>
<li><strong>PFC:</strong> Temporarily pauses the affected traffic class during congestion instead of dropping packets</li>
<li><strong>ECN and CNP:</strong> Detect congestion early and signal the sender to reduce its transmission rate</li>
</ul>
<p>PFC was enabled only for TC3, the traffic class carrying RDMA traffic. SSH and ordinary TCP traffic were not placed in the lossless queue.</p>
</div><div class="fusion-text fusion-text-197"><h3>Why the MikroTik CRS812?</h3>
<p>NVIDIA Spectrum switches are a well-established option for high-speed AI fabrics. This project, however, targeted a compact and cost-conscious four-node lab cluster.</p>
<p>The CRS812 provides the features required for this design, including 400G QSFP-DD ports, breakout support, hardware-offloaded QoS, PFC, ECN, and DCBX. With the correct configuration, approximately 100–111 Gbps of RDMA throughput per link was achieved.</p>
<p>This does not mean that the CRS812 can replace a data-center-class switch in every deployment. For this four-node DGX Spark cluster, however, it provides a practical balance between cost and performance.</p>
</div><div class="fusion-text fusion-text-198"><h3>Preparing the Nodes</h3>
<p>Before cluster setup, all four Spark systems were brought to the same operating system and firmware level. This step is more important than it may appear. Outdated firmware on high-speed network adapters can prevent the expected performance from being reached even when the network configuration is otherwise correct.</p>
<p>The operating system packages and firmware were updated, and the DGX Dashboard was checked for any remaining updates. On the Docker side, all nodes were configured to use the same storage backend, and overlayfs through the containerd snapshotter was verified.</p>
<p>For the management network, the 10GbE port on each Spark was connected to the CRS312 and assigned a static IP address. SSH access, system updates, sparkrun management, and NAS connectivity were all provided through this network.</p>
</div><div class="fusion-text fusion-text-199"><h3>What Does sparkrun Simplify?</h3>
<p>Although four systems can be managed individually, repeatedly running the same commands on every node, configuring SSH keys, and manually launching distributed workloads quickly becomes error-prone.</p>
<p>sparkrun is a toolkit that creates an SSH mesh between DGX Spark nodes and simplifies multi-node model execution. In this deployment, sparkrun was used to:</p>
<ul>
<li>Configure passwordless SSH access between nodes</li>
<li>Prepare the ConnectX-7 networks</li>
<li>Configure user and Docker permissions</li>
<li>Launch the model from a single recipe file</li>
</ul>
<p>Although sparkrun automates many of the setup steps, the host-side DCB configuration was applied separately. DSCP 26 traffic was mapped to priority 3 on the ConnectX-7 interfaces, and PFC was enabled for that priority. A systemd service was created to persist these settings across reboots.</p>
</div><div class="fusion-text fusion-text-200"><h3>Validating the Network</h3>
<p>The network was validated layer by layer before any model workload was launched.</p>
<p>Connectivity between the compute IP addresses was tested first. The <em><strong><code>ping -M do -s 8972</code></strong></em> command was then used to confirm that MTU 9000 worked end to end.</p>
<p>At the TCP layer, <code>iperf3</code> measured approximately 100–120 Gbps of throughput. The RDMA tests produced the following results on each ConnectX-7 subnet:</p>
<ul>
<li>100–111 Gbps write bandwidth</li>
<li>95–110 Gbps read bandwidth</li>
</ul>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-43 hover-type-none"><img decoding="async" width="802" height="456" title="37-ib-write-bw-1" src="https://blog.openzeka.com/en/wp-content/uploads/2026/07/37-ib-write-bw-1.webp" alt class="img-responsive wp-image-1719" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/07/37-ib-write-bw-1-200x114.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/37-ib-write-bw-1-300x171.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/37-ib-write-bw-1-400x227.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/37-ib-write-bw-1-600x341.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/37-ib-write-bw-1-768x437.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/37-ib-write-bw-1-800x455.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/37-ib-write-bw-1.webp 802w" sizes="(max-width: 640px) 100vw, 802px" /></span></div><div class="fusion-text fusion-text-201"><p>&nbsp;</p>
<p>Across both subnets, the Spark systems achieved approximately <strong>200 Gbps of aggregate RDMA bandwidth.</strong> The <strong>ib_write_lat test</strong> measured latency at approximately <strong>1–3 microseconds.</strong></p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-44 hover-type-none"><img decoding="async" width="1024" height="371" title="41-ib-write-lat" src="https://blog.openzeka.com/en/wp-content/uploads/2026/07/41-ib-write-lat-1024x371.webp" alt class="img-responsive wp-image-1720" srcset="https://blog.openzeka.com/en/wp-content/uploads/2026/07/41-ib-write-lat-200x73.webp 200w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/41-ib-write-lat-300x109.webp 300w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/41-ib-write-lat-400x145.webp 400w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/41-ib-write-lat-600x218.webp 600w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/41-ib-write-lat-768x279.webp 768w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/41-ib-write-lat-800x290.webp 800w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/41-ib-write-lat-1024x371.webp 1024w, https://blog.openzeka.com/en/wp-content/uploads/2026/07/41-ib-write-lat.webp 1100w" sizes="(max-width: 640px) 100vw, 1024px" /></span></div><div class="fusion-text fusion-text-202"><p>&nbsp;</p>
<p>These figures matter beyond network benchmarking. When tensor or pipeline parallelism is used, the time required to transfer data between nodes directly affects overall model performance.</p>
</div><div class="fusion-text fusion-text-203"><h3>Running a Real Workload with GLM-5.2</h3>
<p>After the network tests were completed, GLM-5.2 was deployed across four Spark systems using a tensor parallel size of 4.</p>
<p>The model was launched through sparkrun using a single recipe:</p>
<blockquote>
<p>sparkrun run glm52-qt-dcp4-4spark.yaml</p>
</blockquote>
<p>In addition to the model definition, the recipe specifies the container image, memory limits, KV cache settings, quantization format, and vLLM runtime parameters. This makes the deployment reproducible without requiring the full command line to be reconstructed each time.</p>
<p>The <code>Application startup complete.</code> message confirmed that all four nodes were participating in the same distributed model deployment.</p>
<p>In a benchmark performed after the service was ready, the GLM-5.2-Int4 model achieved 20.16 output tokens per second at concurrency 1. This result confirmed that the distributed system could serve a real end-to-end inference request, rather than merely loading the model successfully.</p>
</div><div class="fusion-text fusion-text-204"><h3>Shared Storage</h3>
<p>The ASUSTOR NAS was used as shared storage so that model files did not need to be copied separately to every node. Its two 10GbE ports were connected to the CRS312 using LACP, and the NFS share was mounted on all Spark nodes.</p>
<p>LACP does not automatically increase a single file transfer to 20 Gbps. Its benefit comes from distributing multiple connections and traffic flows across the two physical links.</p>
<p>Model weights, datasets, and benchmark outputs were stored in the shared NAS location.</p>
</div><div class="fusion-text fusion-text-205"><h3>Conclusion</h3>
<p>This project transformed four independent DGX Spark systems into an integrated AI cluster running on shared networking and storage infrastructure. Separating management and compute traffic, implementing RoCEv2/RDMA end to end, centralizing node management with sparkrun, and using a shared NFS volume resulted in an infrastructure that is both manageable and stable.</p>
<p>The resulting platform can run models that exceed the memory capacity of a single Spark. Even when a model already fits on one system, multiple Spark nodes can be used with the right parallelism and serving configuration to increase aggregate token throughput and support more concurrent requests.</p>
<p>The approximately 200 Gbps of aggregate RDMA bandwidth, 1–3 microsecond latency, and successful four-node GLM-5.2 deployment demonstrate that the architecture works not only in theory, but under a real large-model inference workload. With the networking, storage, and cluster-management layers designed as a single system, DGX Spark can serve as the foundation for a compact yet capable distributed AI platform.</p>
<p>All commands, MikroTik switch settings, DCB configuration, network tests, NAS setup, and troubleshooting steps are available in the detailed <a style="color: #00c600;" href="https://whitepapers.openzeka.com/tr/papers/dgx-spark-4node-cluster-kurulumu/"><b>DGX Spark 4-Node AI Cluster Deployment Guide.</b></a></p>
<p>To deploy a similar DGX Spark cluster, the complete system can be ordered from <strong><a style="color: #00c600;" href="https://openzeka.com/en/">openzeka.com</a></strong><span style="color: #13c400;">.</span> Cluster networking, node configuration, and basic commissioning are included at no additional cost with the system purchase.</p>
</div></div></div></div></div>
<p>The post <a href="https://blog.openzeka.com/en/building-a-high-speed-ai-cluster-with-4-dgx-spark-systems/">Building a High-Speed AI Cluster with 4 DGX Spark Systems</a> appeared first on <a href="https://blog.openzeka.com/en">OpenZeka EN Blog</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
