Skip to main content
Toolmodel-servingv1.0

NVIDIA Dynamo

by NVIDIA · open-source · Last verified 2026-04-24

NVIDIA Dynamo is a distributed inference framework for disaggregated prefill and decode serving of large language models across GPU clusters. It enables efficient scaling of inference beyond single-node constraints, distributing KV-cache and attention computation to maximize throughput for very large models on multi-node deployments.

https://github.com/ai-dynamo/dynamo
D
DPoor
Adoption: C+Quality: B+Freshness: ACitations: FEngagement: F

Specifications

License
Open Source
Pricing
open-source
Capabilities
Integrations
Use Cases
API Available
No
SDK Languages
Tags
inference, nvidia, distributed, disaggregated, kv-cache, multi-node, scale
Added
2026-04-24
Completeness
73%

Index Score

34
Adoption
50
Quality
70
Freshness
80
Citations
0
Engagement
0

Need this tool deployed for your team?

Get a Custom Setup

Explore the full AI ecosystem on Agents as a Service