{"id":717,"date":"2025-07-08T13:35:51","date_gmt":"2025-07-08T13:35:51","guid":{"rendered":"https:\/\/blog.aetherix.com\/?p=717"},"modified":"2026-03-27T13:13:18","modified_gmt":"2026-03-27T13:13:18","slug":"jetson-generative-ai-llamaspeak","status":"publish","type":"post","link":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/","title":{"rendered":"Jetson Generative AI \u2013 LlamaSpeak"},"content":{"rendered":"<div class=\"fusion-fullwidth fullwidth-box fusion-builder-row-1 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling\" style=\"--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-padding-right:0px;--awb-padding-left:0px;--awb-flex-wrap:wrap;\" ><div class=\"fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap\" style=\"max-width:1331.2px;margin-left: calc(-4% \/ 2 );margin-right: calc(-4% \/ 2 );\"><div class=\"fusion-layout-column fusion_builder_column fusion-builder-column-0 fusion_builder_column_1_1 1_1 fusion-flex-column\" style=\"--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;\"><div class=\"fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column\"><div class=\"fusion-text fusion-text-1 fusion-text-no-margin\"><p>Vision wasn&#8217;t the only modality transformed by Transformers\u2014**LlamaSpeak** brings the power of large language models to spoken conversations, streaming speech-to-text (ASR), intelligent response generation, and text-to-speech (TTS) back out <strong>in real time<\/strong>\u00a0on your Jetson device.<\/p>\n<p>In this article you&#8217;ll learn how to run NanoLLM&#8217;s <strong>WebChat agent<\/strong> (nicknamed <strong>LlamaSpeak<\/strong>) on Jetson using NVIDIA TensorRT\/MLC and Riva Speech Skills.<\/p>\n<\/div><div class=\"fusion-title title fusion-title-1 fusion-sep-none fusion-title-text fusion-title-size-two\" style=\"--awb-font-size:28px;\"><h2 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\"><h3>Requirements<\/h3><\/h2><\/div>\n<div class=\"table-1\">\n<table width=\"100%\">\n<thead>\n<tr>\n<th align=\"left\">\n<div>\n<div>Hardware \/ Software<\/div>\n<\/div>\n<\/th>\n<th align=\"left\">\n<div>\n<div>Notes<\/div>\n<\/div>\n<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td align=\"left\">\n<div>\n<div><strong>Jetson AI Kit \/ Dev Kit<\/strong><\/div>\n<\/div>\n<\/td>\n<td align=\"left\">\n<div>\n<div>Orin AGX \/ Orin NX recommended for best latency<\/div>\n<\/div>\n<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">\n<div>\n<div><strong>JetPack 6 (L4T r36.x)<\/strong><\/div>\n<\/div>\n<\/td>\n<td align=\"left\">\n<div>\n<div>Needed for latest pre-built containers<\/div>\n<\/div>\n<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">\n<div>\n<div><strong>USB microphone &amp; speakers \/ headset<\/strong><\/div>\n<\/div>\n<\/td>\n<td align=\"left\">\n<div>\n<div>Confirm with `arecord -l`<\/div>\n<\/div>\n<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">\n<div>\n<div><strong>Riva Speech Skills 2.15+<\/strong><\/div>\n<\/div>\n<\/td>\n<td align=\"left\">\n<div>\n<div>Provides low-latency ASR engine<\/div>\n<\/div>\n<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">\n<div>\n<div><strong>Hugging Face token<\/strong><\/div>\n<\/div>\n<\/td>\n<td align=\"left\">\n<div>\n<div>Needed for gated Meta-Llama weights<\/div>\n<\/div>\n<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<div class=\"fusion-title title fusion-title-2 fusion-sep-none fusion-title-text fusion-title-size-two\" style=\"--awb-margin-top:50px;--awb-font-size:28px;\"><h2 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">Obtaining Your Hugging Face Token<\/h2><\/div><div class=\"fusion-text fusion-text-2 fusion-text-no-margin\" style=\"--awb-margin-bottom:10px;\"><p>To download the gated Llama checkpoints you&#8217;ll need a personal access token (PAT) from Hugging Face:<\/p>\n<div>1. Create \/ sign into your account at https:\/\/huggingface.co<\/div>\n<div>2. Click your avatar \u25b8 <strong>Settings<\/strong> \u25b8 <strong>Access Tokens<\/strong><\/div>\n<div>3. Press <strong>New Token<\/strong>, select _&#8221;Read&#8221;_ scope, name it (e.g., `jetson-llamaspeak`), and click <strong>Generate<\/strong><\/div>\n<div>4. Copy the token string that starts with `hf_` \u2014 you&#8217;ll export it in Step 7 below:<\/div>\n<\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-1 > .CodeMirror, .fusion-syntax-highlighter-1 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-1 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_1\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_1\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_1\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\">export HUGGINGFACE_TOKEN=hf_yourlongstringhere<\/textarea><\/div><div class=\"fusion-alert alert notice alert-warning fusion-alert-center awb-alert-native-link-color alert-dismissable awb-alert-close-boxed\" style=\"--awb-margin-top:25px;\" role=\"alert\"><div class=\"fusion-alert-content-wrapper\"><span class=\"alert-icon\"><i class=\"awb-icon-cog\" aria-hidden=\"true\"><\/i><\/span><span class=\"fusion-alert-content\"><em><strong>Tip :<\/strong>\u00a0keep this token private; treat it like a password.<\/em><\/span><\/div><button type=\"button\" class=\"close toggle-alert\" data-dismiss=\"alert\" aria-label=\"Close\">&times;<\/button><\/div><div class=\"fusion-title title fusion-title-3 fusion-sep-none fusion-title-text fusion-title-size-two\" style=\"--awb-font-size:28px;\"><h2 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">Step-by-Step Setup<\/h2><\/div><div class=\"fusion-title title fusion-title-4 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">1.\u00a0 Clone the repository<\/h3><\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-2 > .CodeMirror, .fusion-syntax-highlighter-2 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-2 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_2\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_2\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_2\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\">git clone https:\/\/github.com\/dusty-nv\/jetson-containers<\/textarea><\/div><div class=\"fusion-title title fusion-title-5 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-top:20px;--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">2.\u00a0 Enter the repo<\/h3><\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-3 > .CodeMirror, .fusion-syntax-highlighter-3 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-3 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_3\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_3\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_3\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\">cd jetson-containers<\/textarea><\/div><div class=\"fusion-title title fusion-title-6 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-top:20px;--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">3.\u00a0 Update APT &amp; install pip3<\/h3><\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-4 > .CodeMirror, .fusion-syntax-highlighter-4 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-4 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_4\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_4\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_4\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\">sudo apt update\nsudo apt install -y python3-pip<\/textarea><\/div><div class=\"fusion-title title fusion-title-7 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-top:20px;--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">4. Install helper Python packages<\/h3><\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-5 > .CodeMirror, .fusion-syntax-highlighter-5 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-5 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_5\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_5\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_5\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\">pip3 install -r requirements.txt<\/textarea><\/div><div class=\"fusion-title title fusion-title-8 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-top:20px;--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">5.\u00a0 Verify audio devices<\/h3><\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-6 > .CodeMirror, .fusion-syntax-highlighter-6 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-6 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_6\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_6\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_6\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\">arecord -l   # list capture devices<\/textarea><\/div><div class=\"fusion-text fusion-text-3 fusion-text-no-margin\" style=\"--awb-margin-top:10px;\"><p>If your microphone doesn&#8217;t appear, check USB connections and reboot.<\/p>\n<\/div><div class=\"fusion-title title fusion-title-9 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-top:20px;--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">6.\u00a0 Start the Riva server<\/h3><\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-7 > .CodeMirror, .fusion-syntax-highlighter-7 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-7 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_7\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_7\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_7\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\"># Download the Riva Quick Start ARM64 package from:\n# https:\/\/catalog.ngc.nvidia.com\/orgs\/nvidia\/teams\/riva\/resources\/riva_quickstart_arm64\n\n# Navigate to the downloaded directory:\ncd riva_quickstart_arm64_v2.19.0\n\n# Initialize the Riva environment:\nbash riva_init.sh\n\n# Start the Riva server:\nbash riva_start.sh\n\n# Wait until the server is fully initialized.\n# Proceed only after you see a message like \"Riva server is ready\".<\/textarea><\/div><div class=\"fusion-text fusion-text-4 fusion-text-no-margin\" style=\"--awb-margin-top:10px;\"><p>Accept the license prompt and wait until the log shows <strong>State = READY<\/strong>.<\/p>\n<\/div><div class=\"fusion-title title fusion-title-10 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-top:20px;--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">7.\u00a0 Launch LlamaSpeak<\/h3><\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-8 > .CodeMirror, .fusion-syntax-highlighter-8 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-8 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_8\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_8\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_8\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\">export HUGGINGFACE_TOKEN=hf_xxxxxxxxxxxxxxxxxxxx\n\njetson-containers run --env HUGGINGFACE_TOKEN \\\n  $(autotag nano_llm) \\\n  python3 -m nano_llm.agents.web_chat --api=mlc \\\n    --model meta-llama\/Meta-Llama-3-8B-Instruct \\\n    --asr=riva --tts=piper<\/textarea><\/div><div class=\"fusion-text fusion-text-5 fusion-text-no-margin\" style=\"--awb-margin-top:10px;\"><p>The first run downloads the model (~ 9 GB for 8-bit, ~4 GB for 4-bit) and container layers.<\/p>\n<\/div><div class=\"fusion-title title fusion-title-11 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-top:20px;--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">8.\u00a0 Open the Web UI<\/h3><\/div><div class=\"fusion-text fusion-text-6\"><div>\n<div>When the console prints:<\/div>\n<\/div>\n<\/div><div class=\"fusion-text fusion-text-7\"><div>\n<blockquote>\n<div>WebChat serving at https:\/\/0.0.0.0:8050<\/div>\n<\/blockquote>\n<\/div>\n<\/div><div class=\"fusion-text fusion-text-8\"><div><strong>On-device:<\/strong>\u00a0open a browser on the Jetson and visit `https:\/\/localhost:8050`<\/div>\n<div><strong>Remote:<\/strong>\u00a0replace `&lt;jetson-ip&gt;` with your board&#8217;s address:<\/div>\n<\/div><div class=\"fusion-text fusion-text-9\"><div>\n<blockquote>\n<div>https:\/\/&lt;jetson-ip&gt;:8050<\/div>\n<\/blockquote>\n<\/div>\n<\/div><div class=\"fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-1\" style=\"text-align:center;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);\"><span class=\" fusion-imageframe imageframe-none imageframe-1 hover-type-none\"><img decoding=\"async\" width=\"1024\" height=\"583\" src=\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/llm_chat_screenshot-1024x583.png\" alt class=\"img-responsive wp-image-722\" srcset=\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/llm_chat_screenshot-200x114.png 200w, https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/llm_chat_screenshot-400x228.png 400w, https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/llm_chat_screenshot-600x341.png 600w, https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/llm_chat_screenshot-800x455.png 800w, https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/llm_chat_screenshot-1200x683.png 1200w\" sizes=\"(max-width: 640px) 100vw, 1024px\" \/><\/span><div class=\"awb-imageframe-caption-container\" style=\"text-align:center;\"><div class=\"awb-imageframe-caption\"><p class=\"awb-imageframe-caption-text\">LLM Chat Interface<\/p><\/div><\/div><\/div><div class=\"fusion-title title fusion-title-12 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-top:20px;--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">9.\u00a0 Talk to your Jetson<\/h3><\/div><div class=\"fusion-text fusion-text-10\"><div>\n<div>Grant microphone access when prompted. Try interrupting mid-reply; LlamaSpeak will pause TTS and listen.<\/div>\n<\/div>\n<\/div><div class=\"fusion-title title fusion-title-13 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-top:20px;--awb-margin-bottom:4px;--awb-font-size:18px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;font-size:1em;\">10.\u00a0 (Advanced) Enable multimodality<\/h3><\/div><div class=\"fusion-text fusion-text-11 fusion-text-no-margin\"><p>Launch it with a VLM and drag an image to the chat:<\/p>\n<\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-9 > .CodeMirror, .fusion-syntax-highlighter-9 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-9 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_9\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_9\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_9\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\">--model Efficient-Large-Model\/VILA-7b<\/textarea><\/div><div class=\"fusion-text fusion-text-12\"><p>&nbsp;<\/p>\n<p>See the NanoVLM docs for more supported checkpoints.<\/p>\n<div><\/div>\n<\/div><div class=\"fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-2\" style=\"text-align:center;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);\"><span class=\" fusion-imageframe imageframe-none imageframe-2 hover-type-none\"><img decoding=\"async\" width=\"1024\" height=\"979\" src=\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/multimodal_chat_screenshot-1024x979.png\" alt class=\"img-responsive wp-image-723\" srcset=\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/multimodal_chat_screenshot-200x191.png 200w, https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/multimodal_chat_screenshot-400x382.png 400w, https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/multimodal_chat_screenshot-600x573.png 600w, https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/multimodal_chat_screenshot-800x765.png 800w, https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/multimodal_chat_screenshot-1200x1147.png 1200w, https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/multimodal_chat_screenshot.png 1222w\" sizes=\"(max-width: 640px) 100vw, 1024px\" \/><\/span><div class=\"awb-imageframe-caption-container\" style=\"text-align:center;\"><div class=\"awb-imageframe-caption\"><p class=\"awb-imageframe-caption-text\">Multimodal Chat Interface<\/p><\/div><\/div><\/div><div class=\"fusion-title title fusion-title-14 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><div>\n<h4>LlamaSpeak in Action<\/h4>\n<\/div><\/h1><\/div><div class=\"fusion-text fusion-text-13\"><div>\n<div>Below are demonstration videos showing LlamaSpeak&#8217;s capabilities in both text-only and multimodal scenarios.<\/div>\n<\/div>\n<\/div><div class=\"fusion-title title fusion-title-15 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><div>\n<h4>Text-Only LLM Chat<\/h4>\n<\/div><\/h1><\/div><div class=\"fusion-text fusion-text-14\"><div>\n<div>Experience natural voice conversations with LlamaSpeak using the Meta-Llama-3-8B-Instruct model. The system provides real-time speech recognition, intelligent responses, and natural text-to-speech output.<\/div>\n<\/div>\n<\/div><div class=\"fusion-title title fusion-title-16 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><div>\n<h4>LLM Demo Video:<\/h4>\n<\/div><\/h1><\/div><div class=\"fusion-video fusion-selfhosted-video\" style=\"max-width:100%;\"><div class=\"video-wrapper\"><video playsinline=\"true\" width=\"100%\" style=\"object-fit: cover;\" autoplay=\"true\" muted=\"true\" loop=\"true\" preload=\"auto\" controls=\"1\"><source src=\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/llamaspeak_llm_demo.mp4\" type=\"video\/mp4\">Sorry, your browser doesn&#039;t support embedded videos.<\/video><\/div><\/div><div class=\"fusion-title title fusion-title-17 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><div>\n<h4>Multimodal Vision Chat<\/h4>\n<\/div><\/h1><\/div><div class=\"fusion-text fusion-text-15\"><div>\n<div>See how LlamaSpeak analyzes images while maintaining voice conversation using the VILA-7b vision-language model. Upload images by dragging them into the chat interface.<\/div>\n<\/div>\n<\/div><div class=\"fusion-title title fusion-title-18 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><div>\n<h4>Multimodal Demo Video:<\/h4>\n<\/div><\/h1><\/div><div class=\"fusion-video fusion-selfhosted-video\" style=\"max-width:100%;\"><div class=\"video-wrapper\"><video playsinline=\"true\" width=\"100%\" style=\"object-fit: cover;\" autoplay=\"true\" muted=\"true\" loop=\"true\" preload=\"auto\" controls=\"1\"><source src=\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/llamaspeak_multimodal_demo.mp4\" type=\"video\/mp4\">Sorry, your browser doesn&#039;t support embedded videos.<\/video><\/div><\/div><div class=\"fusion-title title fusion-title-19 fusion-sep-none fusion-title-text fusion-title-size-three\" style=\"--awb-margin-bottom:-10px;\"><h3 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><h3>Troubleshooting<\/h3><\/h3><\/div><div class=\"fusion-title title fusion-title-20 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><h4>How to resolve the NGC Registry access 401 Unauthorized error?<\/h4><\/h1><\/div><div class=\"fusion-text fusion-text-16\"><p>After signing up and logging in to the NGC Catalog website, an API key must be generated from the <strong>&#8220;Setup&#8221;<\/strong> section. Then, run the command <strong><code data-start=\"643\" data-end=\"665\"><span style=\"color: #141617;\">docker login nvcr.io<\/span><\/code><\/strong> in the terminal. In the login prompt, enter <code data-start=\"710\" data-end=\"723\">$oauthtoken<\/code> as the username and the generated API key as the password. If the message<strong> &#8220;Login Succeeded&#8221;<\/strong> appears, the process has been successfully completed.<\/p>\n<\/div><div class=\"fusion-title title fusion-title-21 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><h4>&#8220;model-repo&#8221; Error When Running Riva Manually Without <code data-start=\"197\" data-end=\"212\">riva_start.sh<\/code><\/h4><\/h1><\/div><div class=\"fusion-text fusion-text-17\"><p data-start=\"231\" data-end=\"463\">The tutorial only invokes <code data-start=\"257\" data-end=\"272\"><span style=\"color: #119900;\">riva_start.sh<\/span><\/code>, which handles model repository setup internally. However, when running Riva manually using <code data-start=\"365\" data-end=\"377\">docker run<\/code>, we encountered a <code data-start=\"396\" data-end=\"410\">\"model-repo\"<\/code> error due to missing model repository configuration.<\/p>\n<p data-end=\"571\" data-start=\"465\">To resolve this, we explicitly defined the model path using both an environment variable and a bind mount:<\/p>\n<\/div><style type=\"text\/css\" scopped=\"scopped\">.fusion-syntax-highlighter-10 > .CodeMirror, .fusion-syntax-highlighter-10 > .CodeMirror .CodeMirror-gutters {background-color:#304148;}<\/style><div class=\"fusion-syntax-highlighter-container fusion-syntax-highlighter-10 fusion-syntax-highlighter-theme-dark\" style=\"opacity:0;margin-top:0px;margin-right:0px;margin-bottom:0px;margin-left:0px;font-size:14px;border-width:1px;border-style:solid;border-color:rgba(242,243,245,0);\"><div class=\"syntax-highlighter-copy-code\"><span class=\"syntax-highlighter-copy-code-title\" data-id=\"fusion_syntax_highlighter_10\" style=\"font-size:14px;\">Copy to Clipboard<\/span><\/div><label for=\"fusion_syntax_highlighter_10\" class=\"screen-reader-text\">Syntax Highlighter<\/label><textarea class=\"fusion-syntax-highlighter-textarea\" id=\"fusion_syntax_highlighter_10\" data-readOnly=\"nocursor\" data-lineNumbers=\"\" data-lineWrapping=\"\" data-theme=\"oceanic-next\" data-mode=\"text\/x-sh\">-e RIVA_MODEL_REPO=\/data\/models \\\n-v ~\/riva_quickstart\/model_repository\/models:\/data\/models<\/textarea><\/div><div class=\"fusion-text fusion-text-18\" style=\"--awb-margin-top:20px;\"><p>By setting the <code data-start=\"693\" data-end=\"710\">RIVA_MODEL_REPO<\/code> environment variable and bind-mounting the local <code data-start=\"760\" data-end=\"768\">models<\/code> directory to the container path <code data-start=\"801\" data-end=\"815\">\/data\/models<\/code>, the error was resolved successfully.<\/p>\n<\/div><div class=\"fusion-title title fusion-title-22 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><h4 data-start=\"86\" data-end=\"156\"><strong data-start=\"90\" data-end=\"156\">How to Prevent Port Conflicts (&#8220;port is already in use&#8221; Error)<\/strong><\/h4>\n<h4 data-start=\"158\" data-end=\"462\"><\/h4><\/h1><\/div><div class=\"fusion-text fusion-text-19\"><p>The default port (e.g., 8050) may already be occupied by another service, leading to a &#8220;port is already in use&#8221; error. To prevent this, specify custom ports when launching Riva by using flags like <code data-start=\"355\" data-end=\"372\"><span style=\"color: #119900;\">--web-port 8443<\/span><\/code> and <code data-start=\"377\" data-end=\"393\"><span style=\"color: #119900;\">--ws-port 9443<\/span><\/code>. This ensures Riva runs without interfering with other applications.<\/p>\n<\/div><div class=\"fusion-title title fusion-title-23 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><h4 data-start=\"107\" data-end=\"163\"><strong data-start=\"111\" data-end=\"163\">How to Resolve the Microphone Not Detected Issue<\/strong><\/h4>\n<h4 data-start=\"165\" data-end=\"260\"><\/h4><\/h1><\/div><div class=\"fusion-text fusion-text-20\"><p>When connecting via HTTPS through the browser, sometimes the browser may not automatically prompt for microphone or speaker access permissions. In such cases, the user manually bypasses the security warning by selecting <strong>&#8220;Advanced&#8221; &gt; &#8220;Proceed.&#8221;<\/strong> After the page loads, microphone and speaker permissions are granted manually by clicking the padlock icon in the address bar.<\/p>\n<\/div><div class=\"fusion-title title fusion-title-24 fusion-sep-none fusion-title-text fusion-title-size-one\"><h1 class=\"fusion-title-heading title-heading-left\" style=\"margin:0;\"><h4>How to Fix the Riva Logs Getting Stuck in an Infinite \u201cWaiting\u201d Loop<\/h4><\/h1><\/div><div class=\"fusion-text fusion-text-21\"><p data-start=\"141\" data-end=\"209\">If the Riva logs show an endless \u201cwaiting\u201d loop, follow these steps:<\/p>\n<ul data-start=\"211\" data-end=\"410\">\n<li data-start=\"211\" data-end=\"276\">\n<p data-start=\"213\" data-end=\"276\">Wait until the status changes to <strong data-start=\"246\" data-end=\"255\">READY<\/strong> before proceeding.<\/p>\n<\/li>\n<li data-start=\"277\" data-end=\"410\">\n<p data-start=\"279\" data-end=\"410\">If the loop persists, it may indicate insufficient VRAM. In that case, try selecting a smaller model to reduce memory requirements.<\/p>\n<\/li>\n<\/ul>\n<p data-start=\"412\" data-end=\"497\" data-is-last-node=\"\" data-is-only-node=\"\">This approach helps resolve the \u201cwaiting\u201d hang and ensures Riva initializes properly.<\/p>\n<\/div>\n<div class=\"table-1\">\n<p>&nbsp;<\/p>\n<table width=\"100%\">\n<thead>\n<tr>\n<th align=\"left\">\n<div>\n<div>Issue<\/div>\n<\/div>\n<\/th>\n<th align=\"left\">\n<div>\n<div>Fix<\/div>\n<\/div>\n<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td align=\"left\">\n<div>\n<div>Mic not detected<\/div>\n<\/div>\n<\/td>\n<td align=\"left\">\n<div>\n<div>Use a powered USB sound card; reconfirm with `arecord -l`<\/div>\n<\/div>\n<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">\n<div>\n<div>Stuck on Waiting for Riva<\/div>\n<\/div>\n<\/td>\n<td align=\"left\">\n<div>\n<div>Ensure the Riva container is running and ports are exposed<\/div>\n<\/div>\n<\/td>\n<\/tr>\n<tr>\n<td align=\"left\">\n<div>\n<div>Out-of-memory errors<\/div>\n<\/div>\n<\/td>\n<td align=\"left\">\n<div>\n<div>Use a 4-bit quantized checkpoint or a 7-B model (e.g., Mistral-7B)<\/div>\n<\/div>\n<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<div class=\"fusion-text fusion-text-22\" style=\"--awb-text-color:#000000;--awb-margin-top:30px;\"><p><em><strong>For more information about NanoLLM and advanced configurations, visit the <a style=\"color: #119900;\" href=\"https:\/\/github.com\/dusty-nv\/NanoLLM\"><span style=\"color: #119900;\">NanoLLM GitHub reposito<\/span>ry<\/a><span style=\"color: #119900;\">.<\/span><\/strong><\/em><\/p>\n<\/div><\/div><\/div><\/div><\/div>\n","protected":false},"excerpt":{"rendered":"","protected":false},"author":3,"featured_media":1577,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[17],"tags":[],"class_list":["post-717","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-generative-ai"],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v25.3.1 (Yoast SEO v25.3.1) - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>LlamaSpeak on Jetson: Real-Time Voice AI<\/title>\n<meta name=\"description\" content=\"Set up LlamaSpeak on Jetson for fast voice AI using Meta-Llama, Riva, and NanoLLM. Simple step-by-step guide.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Jetson Generative AI \u2013 LlamaSpeak\" \/>\n<meta property=\"og:description\" content=\"Set up LlamaSpeak on Jetson for fast voice AI using Meta-Llama, Riva, and NanoLLM. Simple step-by-step guide.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/\" \/>\n<meta property=\"og:site_name\" content=\"OpenZeka EN Blog\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/profile.php?id=61576911356211\" \/>\n<meta property=\"article:published_time\" content=\"2025-07-08T13:35:51+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-03-27T13:13:18+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1920\" \/>\n\t<meta property=\"og:image:height\" content=\"1500\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Enhar\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@Aetherixnl\" \/>\n<meta name=\"twitter:site\" content=\"@Aetherixnl\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Enhar\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"5 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/\"},\"author\":{\"name\":\"Enhar\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/#\/schema\/person\/62c964376839cf2c4b2eb682bf14d3cb\"},\"headline\":\"Jetson Generative AI \u2013 LlamaSpeak\",\"datePublished\":\"2025-07-08T13:35:51+00:00\",\"dateModified\":\"2026-03-27T13:13:18+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/\"},\"wordCount\":4834,\"publisher\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/#organization\"},\"image\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#primaryimage\"},\"thumbnailUrl\":\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg\",\"articleSection\":[\"Generative AI\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/\",\"url\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/\",\"name\":\"LlamaSpeak on Jetson: Real-Time Voice AI\",\"isPartOf\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#primaryimage\"},\"image\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#primaryimage\"},\"thumbnailUrl\":\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg\",\"datePublished\":\"2025-07-08T13:35:51+00:00\",\"dateModified\":\"2026-03-27T13:13:18+00:00\",\"description\":\"Set up LlamaSpeak on Jetson for fast voice AI using Meta-Llama, Riva, and NanoLLM. Simple step-by-step guide.\",\"breadcrumb\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#primaryimage\",\"url\":\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg\",\"contentUrl\":\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg\",\"width\":1920,\"height\":1500},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/blog.openzeka.com\/en\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Jetson Generative AI \u2013 LlamaSpeak\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/#website\",\"url\":\"https:\/\/blog.openzeka.com\/en\/\",\"name\":\"Aetherix B.V.\",\"description\":\"NVIDIA Jetson Developer Kits &amp;Edge Devices\",\"publisher\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/blog.openzeka.com\/en\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/#organization\",\"name\":\"Aetherix B.V.\",\"url\":\"https:\/\/blog.openzeka.com\/en\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/06\/aetherix-site-icon.webp\",\"contentUrl\":\"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/06\/aetherix-site-icon.webp\",\"width\":421,\"height\":398,\"caption\":\"Aetherix B.V.\"},\"image\":{\"@id\":\"https:\/\/blog.openzeka.com\/en\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/profile.php?id=61576911356211\",\"https:\/\/x.com\/Aetherixnl\",\"https:\/\/www.instagram.com\/aetherixnl\/\",\"https:\/\/www.tiktok.com\/@aetherixnl\"],\"description\":\"Aetherix provides a full range of NVIDIA Jetson-based edge AI solutions\u2014including Developer Kits, AI Kits, industrial-grade Carrier Boards, and fully integrated Boxed AI Systems.\",\"email\":\"info@aetherix.com\",\"legalName\":\"Aetherix B.V.\",\"vatID\":\"NL867727688B01\"},{\"@type\":\"Person\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/#\/schema\/person\/62c964376839cf2c4b2eb682bf14d3cb\",\"name\":\"Enhar\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/blog.openzeka.com\/en\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/189d567adce3bb0c8d438b4586bf861ec04980f2e451003975e3cf871781d0f4?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/189d567adce3bb0c8d438b4586bf861ec04980f2e451003975e3cf871781d0f4?s=96&d=mm&r=g\",\"caption\":\"Enhar\"}}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"LlamaSpeak on Jetson: Real-Time Voice AI","description":"Set up LlamaSpeak on Jetson for fast voice AI using Meta-Llama, Riva, and NanoLLM. Simple step-by-step guide.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/","og_locale":"en_US","og_type":"article","og_title":"Jetson Generative AI \u2013 LlamaSpeak","og_description":"Set up LlamaSpeak on Jetson for fast voice AI using Meta-Llama, Riva, and NanoLLM. Simple step-by-step guide.","og_url":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/","og_site_name":"OpenZeka EN Blog","article_publisher":"https:\/\/www.facebook.com\/profile.php?id=61576911356211","article_published_time":"2025-07-08T13:35:51+00:00","article_modified_time":"2026-03-27T13:13:18+00:00","og_image":[{"width":1920,"height":1500,"url":"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg","type":"image\/jpeg"}],"author":"Enhar","twitter_card":"summary_large_image","twitter_creator":"@Aetherixnl","twitter_site":"@Aetherixnl","twitter_misc":{"Written by":"Enhar","Est. reading time":"5 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#article","isPartOf":{"@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/"},"author":{"name":"Enhar","@id":"https:\/\/blog.openzeka.com\/en\/#\/schema\/person\/62c964376839cf2c4b2eb682bf14d3cb"},"headline":"Jetson Generative AI \u2013 LlamaSpeak","datePublished":"2025-07-08T13:35:51+00:00","dateModified":"2026-03-27T13:13:18+00:00","mainEntityOfPage":{"@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/"},"wordCount":4834,"publisher":{"@id":"https:\/\/blog.openzeka.com\/en\/#organization"},"image":{"@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#primaryimage"},"thumbnailUrl":"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg","articleSection":["Generative AI"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/","url":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/","name":"LlamaSpeak on Jetson: Real-Time Voice AI","isPartOf":{"@id":"https:\/\/blog.openzeka.com\/en\/#website"},"primaryImageOfPage":{"@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#primaryimage"},"image":{"@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#primaryimage"},"thumbnailUrl":"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg","datePublished":"2025-07-08T13:35:51+00:00","dateModified":"2026-03-27T13:13:18+00:00","description":"Set up LlamaSpeak on Jetson for fast voice AI using Meta-Llama, Riva, and NanoLLM. Simple step-by-step guide.","breadcrumb":{"@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#primaryimage","url":"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg","contentUrl":"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/07\/9.jpg","width":1920,"height":1500},{"@type":"BreadcrumbList","@id":"https:\/\/blog.openzeka.com\/en\/jetson-generative-ai-llamaspeak\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/blog.openzeka.com\/en\/"},{"@type":"ListItem","position":2,"name":"Jetson Generative AI \u2013 LlamaSpeak"}]},{"@type":"WebSite","@id":"https:\/\/blog.openzeka.com\/en\/#website","url":"https:\/\/blog.openzeka.com\/en\/","name":"Aetherix B.V.","description":"NVIDIA Jetson Developer Kits &amp;Edge Devices","publisher":{"@id":"https:\/\/blog.openzeka.com\/en\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/blog.openzeka.com\/en\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/blog.openzeka.com\/en\/#organization","name":"Aetherix B.V.","url":"https:\/\/blog.openzeka.com\/en\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/blog.openzeka.com\/en\/#\/schema\/logo\/image\/","url":"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/06\/aetherix-site-icon.webp","contentUrl":"https:\/\/blog.openzeka.com\/en\/wp-content\/uploads\/2025\/06\/aetherix-site-icon.webp","width":421,"height":398,"caption":"Aetherix B.V."},"image":{"@id":"https:\/\/blog.openzeka.com\/en\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/profile.php?id=61576911356211","https:\/\/x.com\/Aetherixnl","https:\/\/www.instagram.com\/aetherixnl\/","https:\/\/www.tiktok.com\/@aetherixnl"],"description":"Aetherix provides a full range of NVIDIA Jetson-based edge AI solutions\u2014including Developer Kits, AI Kits, industrial-grade Carrier Boards, and fully integrated Boxed AI Systems.","email":"info@aetherix.com","legalName":"Aetherix B.V.","vatID":"NL867727688B01"},{"@type":"Person","@id":"https:\/\/blog.openzeka.com\/en\/#\/schema\/person\/62c964376839cf2c4b2eb682bf14d3cb","name":"Enhar","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/blog.openzeka.com\/en\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/189d567adce3bb0c8d438b4586bf861ec04980f2e451003975e3cf871781d0f4?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/189d567adce3bb0c8d438b4586bf861ec04980f2e451003975e3cf871781d0f4?s=96&d=mm&r=g","caption":"Enhar"}}]}},"_links":{"self":[{"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/posts\/717","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/comments?post=717"}],"version-history":[{"count":29,"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/posts\/717\/revisions"}],"predecessor-version":[{"id":1555,"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/posts\/717\/revisions\/1555"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/media\/1577"}],"wp:attachment":[{"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/media?parent=717"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/categories?post=717"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.openzeka.com\/en\/wp-json\/wp\/v2\/tags?post=717"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}