aoniAI Hub

支持的模型

在 Jetson 设备上运行最新的 AI 大模型

硬件:
引擎:

36 个模型

Gemma 3(5)

LLM4B

Gemma 3 4B

Google Gemma 3 系列 4B 语言模型

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

Google

LLM1B

Gemma 3 1B

Google Gemma 3 系列 1B 轻量语言模型,适合资源受限设备

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

Google

NEWLLM270M

Gemma 3 270M

Google Gemma 3 系列 270M 超轻量语言模型

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

Google

LLM27B

Gemma 3 27B

Google Gemma 3 系列 27B 大型语言模型

thor 128gbthor 64gborin 64gb
引擎vllm

Google

LLM12B

Gemma 3 12B

Google Gemma 3 系列 12B 中型语言模型

thor 128gbthor 64gborin 64gborin 16gb
引擎vllm

Google

Qwen3(5)

LLM8B5.5 GB

Qwen3-VL-8B-Instruct

Qwen3-VL-8B-Instruct 是阿里云通义千问团队开发的一款80亿参数的多模态大模型,在同系列中兼具强大的性能和更友好的部署成本

thor 128gb
引擎vllmTensorRT Edge-LLM

Alibaba

LLM30B-A3B

Qwen3 30B-A3B

阿里巴巴 Qwen3 系列 MoE 模型,30B 总参数仅激活 3B,兼顾性能与效率

thor 128gbthor 64gborin 64gb
引擎vllm

Alibaba

LLM8B5.5 GB

Qwen3-8B

阿里巴巴 Qwen3 系列的中型语言模型,8B 参数,原生支持思考模式,适合单 GPU 部署的通用文本任务

thor 128gbthor 64gborin 64gborin 16gb
引擎vllm

Alibaba

LLM4B

Qwen3 4B

阿里巴巴 Qwen3 系列的小型语言模型,4B 参数,适合资源受限的边缘设备部署

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

Alibaba

LLM32B

Qwen3 32B

阿里巴巴 Qwen3 系列 32B 稠密模型,旗舰级通用能力,原生支持思考模式

thor 128gbthor 64gb
引擎vllm

Alibaba

Gemma 4(2)

LLM26B-A4B

Gemma 4 26B-A4B

Google Gemma 4 系列 MoE 模型,25.8B 总参数 / 3.8B 激活,256K 上下文,支持文本和图像理解

thor 128gbthor 64gborin 64gb
引擎vllm

Google

LLM31B

Gemma 4 31B

Google Gemma 4 系列 30.7B 稠密旗舰模型,256K 上下文,Arena AI 文本排行榜开源模型第 3 名

thor 128gbthor 64gborin 64gb
引擎vllm

Google

Qwen(1)

NEWLLM27B17GB

Qwen3.8-27B-GGUF

Unsloth 团队发布的 Qwen3.8-27B 量化版(GGUF 格式)模型卡,Apache-2.0 许可,约 1.7k 点赞。模型基于 Qwen3.5 架构演进,是开源 Qwen 家族目前最强一代:27B 参数的稠密因果语言模型,配备视觉编码器,原生支持图文与视频理解。亮点包括:灵活思维控制(思考模式默认开启、可按请求关闭,支持 reasoning_effort 调节推理深度);更强的智能体规划与环境反馈处理能力;改进的工具调用解析;兼容 Codex 等智能体工具。存储约 16.3GB,提供多种量化档位(Q3_K、Q4_K_M、Q5_K_M、IQ4_NL、UD-Q4_K_XL 等)及 mmproj 视觉投影。

thor 128gbthor 64gborin 64gb
引擎llama.cpp

Alibaba

Nemotron(3)

VLMVLM30B-A3B

Nemotron-3-Nano-Omni

NVIDIA 推出的全模态 MoE 推理模型,30B 总参数仅激活 3B,原生支持文本、图像、音频、视频四种输入,256K 上下文

thor 128gbthor 64gborin 64gb
引擎vllm

NVIDIA

LLM30B-A3B

Nemotron3 Nano 30B-A3B

NVIDIA Nemotron3 Nano MoE 模型,30B 总参数仅激活 3B

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

NVIDIA

LLM9B

Nemotron Nano 9B v2

NVIDIA Nemotron Nano 9B v2 语言模型

thor 128gbthor 64gb
引擎vllm

NVIDIA

Ministral 3(6)

VLMVLM14B

Ministral 3 14B Instruct

Mistral AI 推出的 14B 指令型视觉语言模型,262K 上下文,FP8 精度

thor 128gbthor 64gborin 64gb
引擎vllm

Mistral AI

VLMVLM14B

Ministral 3 14B Reasoning

Mistral AI 推出的 14B 推理型视觉语言模型,262K 上下文,FP16 精度

thor 128gbthor 64gborin 64gb
引擎vllm

Mistral AI

VLMVLM8B

Ministral 3 8B Reasoning

Mistral AI 推出的 8B 推理型视觉语言模型,262K 上下文,FP16 精度

thor 128gbthor 64gborin 64gborin 16gb
引擎vllm

Mistral AI

VLMVLM3B

Ministral 3 3B Reasoning

Mistral AI 推出的 3B 推理型视觉语言模型,262K 上下文,FP16 精度,专注逻辑推理和问题求解

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

Mistral AI

VLMVLM8B

Ministral 3 8B Instruct

Mistral AI 推出的 8B 指令型视觉语言模型,262K 上下文,FP8 精度

thor 128gbthor 64gborin 64gborin 16gb
引擎vllm

Mistral AI

VLMVLM3B

Ministral 3 3B Instruct

Mistral AI 推出的 3B 指令型视觉语言模型,262K 上下文,FP8 精度,支持多语言、函数调用和视觉理解

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

Mistral AI

Gemma(1)

LLM270M

FunctionGemma

Google 推出的函数调用专用 Gemma 模型,270M 参数,专为工具调用和结构化输出优化

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

Google

Qwen3.5(5)

VLMVLM9B

Qwen3.5 9B

阿里 Qwen3.5 系列 9B 视觉语言模型,262K 上下文,支持图文理解、工具调用和代理

thor 128gbthor 64gborin 64gborin 16gb
引擎vllm

Alibaba

VLMVLM4B

Qwen3.5 4B

阿里 Qwen3.5 系列 4B 视觉语言模型,262K 上下文,AWQ 4bit 量化适合 Jetson Orin 部署

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllmTensorRT Edge-LLM

Alibaba

VLMVLM35B-A3B

Qwen3.5 35B-A3B

阿里 Qwen3.5 系列 MoE 视觉语言模型,总参数 35B 仅激活 3B,262K 上下文,支持图文理解、函数调用和多语言

thor 128gbthor 64gborin 64gb
引擎vllm

Alibaba

VLMVLM27B

Qwen3.5 27B

阿里 Qwen3.5 系列 27B 稠密视觉语言模型,262K 上下文,支持图文理解、推理、函数调用和多语言

thor 128gbthor 64gborin 64gb
引擎vllm

Alibaba

VLMVLM0.8B

Qwen3.5 0.8B

阿里 Qwen3.5 系列最小视觉语言模型,0.8B 参数,262K 上下文,BF16 精度,适合快速原型和轻量边缘部署

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllmTensorRT Edge-LLM

Alibaba

Llama 3.2(1)

LLM3B

Llama 3.2 3B

Meta Llama 3.2 系列 3B 小型语言模型,专为边缘和移动设备优化

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

Meta

GPT OSS(2)

LLM120B

GPT OSS 120B

OpenAI 开源 GPT OSS 120B 大语言模型,NVFP4 精度优化

thor 128gbthor 64gb
引擎vllm

OpenAI

LLM20B

GPT OSS 20B

OpenAI 开源 GPT OSS 20B 模型,NVFP4 精度优化

thor 128gbthor 64gborin 64gb
引擎vllm

OpenAI

Llama 3.1(1)

LLM8B

Llama 3.1 8B

Meta Llama 3.1 系列 8B 通用语言模型,多语言支持和强大的指令遵循能力

thor 128gbthor 64gborin 64gborin 16gb
引擎vllm

Meta

Qwen3.6(2)

LLM27B19 GB

Qwen3.6 27B

阿里 Qwen3.6 系列 27B 稠密模型,19GB NVFP4 量化,支持 MTP 推测解码,强推理和函数调用能力

thor 128gbthor 64gb
引擎vllmTensorRT Edge-LLM

Alibaba

NEWLLM35B-A3B

Qwen3.6 35B-A3B(MoE)

阿里 Qwen3.6 系列 MoE 模型,总参数 35B 仅激活 3B,支持 MTP 推测解码,原生支持推理和函数调用

thor 128gbthor 64gborin 64gb
引擎vllmTensorRT Edge-LLM

Alibaba

Qwen3 VL(2)

VLMVLM4B

Qwen3 VL 4B

阿里巴巴 Qwen3 VL 系列 4B 视觉语言模型,适合边缘设备视觉任务

thor 128gbthor 64gborin 64gborin 16gborin 8gb
引擎vllm

Alibaba

VLMVLM8B

Qwen3 VL 8B

阿里巴巴 Qwen3 VL 系列 8B 视觉语言模型,支持图像理解和文本生成

thor 128gbthor 64gborin 64gborin 16gb
引擎vllm

Alibaba