Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
Inference Optimization
/
Quantization
Quantization
Subtopic
29 questions
Questions tagged with Quantization — part of Inference Optimization.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Deploying a vision-language model on a phone or AR headset instead of a server: what actually changes?
Short Answer
Medium
Match each vector compression scheme to its core mechanism.
Match Pairs
Medium
Match each GPU generation to the quantization format that lights up its native tensor cores.
Match Pairs
Medium
Which quantization combo squeezes the most decode bandwidth per percent of quality lost?
Multiple Choice
Medium
Rank these chat API cost levers from highest to lowest typical ROI
Order Steps
Medium
Match Hopper FP8 and Blackwell FP4 tensor cores to their throughput and bandwidth effects on decode
Match Pairs
Medium
Weight quantization in plain terms: what changes and what stays the same?
Flashcard
Easy
Where does FP8 fit in modern LLM serving and which GPUs support it?
Flashcard
Easy
Match each 4-bit format to its niche in 2026 inference and fine-tuning.
Match Pairs
Easy
Define activation quantization and explain how W8A8 differs from weight only quant.
Flashcard
Easy
Merging a LoRA back into an NF4 quantized base hits one subtle problem: name it and explain why
Short Answer
Hard
FP4 vs NF4, pick the answer that captures the actual structural difference
Multiple Choice
Easy
INT8 KV cache vs INT8 weight quantization, which one is easier in production?
Multiple Choice
Medium
Name two recovery moves a serving stack can pull when KV memory runs short mid request.
Short Answer
Medium
How does QLoRA fit a 65B fine-tune onto a single 48GB GPU?
Short Answer
Hard
Which statements about QLoRA are true?
Multi-select
Medium
Premium
Why does QLoRA use…
Short Answer
Hard
Premium
Misconception: 'LoRA is just…
Multiple Choice
Easy
Which statements about FP8, INT8 and INT4 weight quantization are correct?
Multi-select
Medium
Match each weight quantization regime to what it buys you and where it breaks
Match Pairs
Hard
Premium
What does W8A8 vs…
Short Answer
Hard
Match each weight/activation quant regime to its mechanism and best fit workload
Match Pairs
Hard
Spot the errors in this tensor core utilization claim
Spot the Error
Hard
On the roofline, which intervention moves an LLM decode workload most directly toward higher peak throughput?
Multiple Choice
Hard
Match each post-training quantization method to what it protects against
Match Pairs
Hard
Showing 1–25 of 29
← Prev
Next →
Lyzr
Ola
C3 Ai
Persistent
Cognizant
Dify
Cognizant
Shield Ai
Contextual Ai
Datarobot
Ltimindtree
NVIDIA
Deloitte
Meesho
Mercor
NVIDIA
Flowise
NVIDIA
Accenture
Contextual Ai
Humanloop
Netflix
Flipkart
Graphcore
Hugging Face
NVIDIA
Descript
NVIDIA
Kpmg
Notion
Bytedance
Flowise
Contextual Ai
Vellum
Neptune Ai
NVIDIA
Comet Ml
Persistent
Oracle
Spotify
Induced Ai
Krutrim
Gnani
Harvey
Anduril
Salesforce
Nykaa
Tata Digital
Airbnb
Dataiku