Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
Fp8
Fp8
Subtopic
12 questions
Questions tagged with Fp8 — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Match each GPU generation to the quantization format that lights up its native tensor cores.
Match Pairs
Medium
At batch 16 and 4k context with FP8 weights, which model still fits on one 80 GB H100?
Multiple Choice
Medium
Match Hopper FP8 and Blackwell FP4 tensor cores to their throughput and bandwidth effects on decode
Match Pairs
Medium
Weight quantization in plain terms: what changes and what stays the same?
Flashcard
Easy
Where does FP8 fit in modern LLM serving and which GPUs support it?
Flashcard
Easy
Spell out TensorRT-LLM and pinpoint its production differentiator
Flashcard
Easy
Same LoRA recipe on H100 vs A100: what precision should each team default to?
Multiple Choice
Medium
Which statements about FP8, INT8 and INT4 weight quantization are correct?
Multi-select
Medium
Match each weight quantization regime to what it buys you and where it breaks
Match Pairs
Hard
Match each post-training quantization method to what it protects against
Match Pairs
Hard
How much does FP8 KV cache help, what does it cost and what is KIVI?
Short Answer
Hard
Premium
An engineer says 'FP8…
Multiple Choice
Hard
Mercor
NVIDIA
Flowise
NVIDIA
Humanloop
Netflix
Hugging Face
NVIDIA
NVIDIA
Qualcomm
Dataiku
NVIDIA
Arize Ai
NVIDIA
Oracle
Spotify
Induced Ai
Krutrim
Capgemini
Niki Ai
Nykaa
Tata Digital
Forethought
NVIDIA