The NVIDIA L40 is the first Ada Lovelace 48 GB datacenter GPU, released ahead of the L40S, combining 48 GB of GDDR6 with Ada Tensor Cores in a PCIe form factor. It was positioned for visualization, rendering, and inference workloads. The L40S later superseded it for pure inference with higher INT8 throughput and better compute balance. The L40 remains a capable option for teams running mixed visualization and inference workloads, with 90 TFLOPS FP16 and 720 INT8 TOPS in a 300W envelope. It handles 30B models at Q4 and 13B models at FP16.
Beyond LLMs
AI Capability Matrix
What AI tasks this GPU can handle — from text generation to image and video creation.
48 GB GDDR6 VRAM864 GB/s memory bandwidth90 TFLOPS FP16 / 720 INT8 TOPSAda Lovelace architecture with FP8 Tensor Core supportPCIe 4.0 x16, 300W TDPSupports NVENC/NVDEC for multimedia workloads alongside AI
Para cargas de trabajo de IA
Fortalezas
48 GB VRAM fits 30B models at Q4 and 13B at FP16 comfortably
FP8 Ada Tensor Cores provide a meaningful step up from Ampere A40 at the same VRAM tier
300W TDP is lower than L40S (350W) — slightly better for dense configurations
Handles mixed rendering and inference workloads for studios or labs needing both
Consideraciones
Superseded by L40S for pure inference — L40S offers higher INT8 TOPS for similar price
No MIG support — cannot partition for multi-tenant isolated inference
GDDR6 bandwidth limits token generation speed compared to HBM-based alternatives
Limited new availability; mostly found in the secondary or certified refurbished market
Architecture
Ada Lovelace
Ada Lovelace is NVIDIA's fourth-generation RTX architecture, manufactured on TSMC's custom 4N process. It introduces 4th-generation Tensor Cores with FP8 support, 3rd-generation ray tracing cores, and the Shader Execution Reordering (SER) engine for improved workload scheduling.
AI Relevance
FP8 Tensor Core operations provide a significant uplift for quantized LLM inference compared to Ampere's FP16-only Tensor Cores. DLSS 3 Frame Generation demonstrates the architecture's AI processing capabilities.
Qwen 3.5 27B matches Chat and keeps a practical fit profile. It is a recent-generation family, which helps on current local SOTA workloads. It fits natively with comfortable headroom. Context coverage stays within the requested workload envelope. Known distribution channels: huggingface, ollama, lm-studio.
Qwen 3.6 27B is a specialized fit for Coding. It is a recent-generation family, which helps on current local SOTA workloads. It fits natively with comfortable headroom. Context coverage stays within the requested workload envelope. Known distribution channels: huggingface, lm-studio.
Qwen 3.6 27B is a specialized fit for Agentic Coding. It is a recent-generation family, which helps on current local SOTA workloads. It fits natively with comfortable headroom. Context coverage stays within the requested workload envelope. Known distribution channels: huggingface, lm-studio.
Devstral Small 2 24B Instruct matches Reasoning and keeps a practical fit profile. It is a recent-generation family, which helps on current local SOTA workloads. It fits natively with comfortable headroom. Context coverage stays within the requested workload envelope. Known distribution channels: huggingface, ollama, lm-studio.
Qwen 3.5 27B matches RAG and keeps a practical fit profile. It is a recent-generation family, which helps on current local SOTA workloads. It fits natively with comfortable headroom. Context coverage stays within the requested workload envelope. Known distribution channels: huggingface, ollama, lm-studio.
Image models estimated at 1024×1024 (28 steps, FP16). Video models estimated at 768×512 (25 frames, 30 steps, FP16). Actual performance varies with runtime and system load.
Multi-GPU scaling
NVIDIA L40 48GB — Up to 2× via PCIe
Scale out with multiple GPUs for larger models. PCIe interconnect with 25% scaling overhead.
Config
Effective memory
Models that fit
Est. bandwidth
1× NVIDIA
48 GB
343/380
864 GB/s
2× NVIDIA
96 GB
356/380
1,296 GB/s
Model counts use default quantization at coding workload settings. Multi-GPU scaling factor: 0.75× per additional GPU.
NVIDIA L40 48GB (48 GB VRAM) can run these top models: Qwen 3.6 35B A3B (score: 98/100), Qwen 3.5 35B A3B (score: 96/100), Qwen3-Coder 30B A3B Instruct (score: 96/100). See the full compatibility list above.
How much VRAM does NVIDIA L40 48GB have for AI?
NVIDIA L40 48GB has 48 GB of VRAM available for AI model inference. This determines which models and quantization levels you can run locally.
Is NVIDIA L40 48GB good for running LLMs locally?
Yes, NVIDIA L40 48GB is excellent for running LLMs locally with top compatibility scores above 80/100.
What is the best model for NVIDIA L40 48GB for coding?
For coding on NVIDIA L40 48GB, we recommend Qwen 3.6 27B. It achieves 24.7 tokens per second with 262K context window. Qwen 3.6 27B is a specialized fit for Coding. It is a recent-generation family, which helps on current local SOTA workloads. It fits natively with comfortable headroom. Context coverage stays within the requested workload envelope. Known distribution channels: huggingface, lm-studio.
Should I upgrade from NVIDIA L40 48GB?
There are 5 upgrade path(s) from NVIDIA L40 48GB: NVIDIA L40 48GB, AMD Instinct MI210 64GB. Upgrading would unlock larger models and faster inference speeds.
Can NVIDIA L40 48GB run Flux for image generation?
Yes, NVIDIA L40 48GB with 48 GB of usable memory can run Flux.1 Dev at FP16 natively. Flux is a 12B parameter diffusion transformer that produces high-quality images. You can also run the Schnell variant for faster generation.
What image and video AI models can I run on NVIDIA L40 48GB?
NVIDIA L40 48GB (48 GB VRAM) can handle various AI generation tasks beyond LLMs. For image generation, SDXL and Stable Diffusion 3.5 run well. Flux.1 Dev also runs natively for state-of-the-art image quality. For video, LTX Video 2.3 can generate short clips. Check the AI Capability Matrix above for detailed compatibility.
Is NVIDIA L40 48GB good for AI image generation?
NVIDIA L40 48GB is excellent for AI image generation. With 48 GB of usable memory, it runs all major diffusion models including Flux.1, SDXL, and Stable Diffusion 3.5 at full precision. You can generate high-resolution images quickly and even handle video generation models.
Can NVIDIA L40 48GB run Qwen 3.5 27B?
Yes, NVIDIA L40 48GB with 48 GB of usable memory can run Qwen 3.5 27B at Q8 (near-lossless, ~28.9 GB) or even FP16 (~55.4 GB) depending on your context needs. This setup provides an excellent experience with this model. Use Ollama or vLLM for best results.
What is the best quantization for AI models on NVIDIA L40 48GB?
With 48 GB VRAM on NVIDIA L40 48GB, use Q8_0 for most models — it is near-lossless and you have the memory for it. For 70B+ models, Q6_K offers excellent quality. Reserve Q4_K_M for 100B+ models or when you need maximum context length.
For local LLMs on NVIDIA L40 48GB, does VRAM matter more than bandwidth?
NVIDIA L40 48GB has enough memory for many local LLMs, but bandwidth still matters a lot for real speed. Once a model fits, a faster-memory GPU can feel significantly better than a slower setup with similar capacity.
How does multi-GPU scale for AI inference on NVIDIA L40 48GB?
NVIDIA L40 48GB supports up to 2× GPU scaling via PCIe. With 2× GPUs, you get 96 GB effective memory with a 0.75× scaling factor per GPU. This enables running models like Qwen 3.5 397B A17B and Devstral 2 123B Instruct that don't fit on a single card.
Is PCIe required for multi-GPU NVIDIA L40 48GB inference?
NVIDIA L40 48GB uses PCIe for multi-GPU communication, which has approximately 25% scaling overhead. For best multi-GPU performance, consider NVLink-equipped variants.
Do I need more PCIe lanes or a workstation motherboard for multi-GPU NVIDIA L40 48GB builds?
Usually yes. If you want to run 2-4× NVIDIA L40 48GB for local AI, the bottleneck often becomes the platform, not the card. Workstation and server boards give you more CPU PCIe lanes, better x16 slot wiring, more spacing between cards, stronger power delivery, and usually more RAM capacity. Consumer x8/x8 layouts can work, but they are a common weak point in multi-GPU builds.