Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
H100
H100
Subtopic
11 questions
Questions tagged with H100 — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
At batch 16 and 4k context with FP8 weights, which model still fits on one 80 GB H100?
Multiple Choice
Medium
Pick the H100 workload where spot pricing is genuinely safe to deploy on
Multiple Choice
Medium
Sketch a Llama 3.1 70B self-host versus hosted API break even on the back of an envelope
Short Answer
Medium
A MIG slice on an H100, what is it and what is it good for?
Flashcard
Easy
Identify the NVIDIA H100: what generation, what memory, what tensor core formats?
Flashcard
Easy
Same LoRA recipe on H100 vs A100: what precision should each team default to?
Multiple Choice
Medium
LoRA vs full FT on 7B + 10k examples: order of magnitude cost comparison
Short Answer
Medium
Apply the roofline model to LLM inference and identify where decode and prefill sit on it.
Short Answer
Hard
Walk through the byte accounting that proves a single batch decode step is bandwidth bound on H100.
Short Answer
Hard
Predict the bandwidth vs compute latency of a single Llama-70B decode step on H100
Predict Output
Hard
Predict the critical batch size where decode crosses from memory bound to compute bound on H100
Predict Output
Hard
NVIDIA
Qualcomm
Mckinsey
Mongodb
Dataiku
NVIDIA
Crewai
Glean
Ey
NVIDIA
Meesho
NVIDIA
Intuit
NVIDIA
Amd
Cresta
Dataiku
LangChain
Airbnb
Graphcore
Decagon
Deloitte