Zenaique
Topics
Practice
Study
Browse
Reference
Pricing
Search…
⌘K
Topics
/
LLM System Design
/
Cost
Cost
Subtopic
28 questions
Questions tagged with Cost — part of LLM System Design.
Premium questions for this topic
Format
Difficulty
Role
Experience
Companies
Sort
Newest
Quality
Difficulty ↑
Difficulty ↓
Questions
Premium
Pick the right classifier…
Multiple Choice
Medium
After a cross-encoder reranker was added on every query, costs doubled: keep the quality, cut the bill
Short Answer
Medium
Predict the prompt token count before and after context compression
Predict Output
Medium
At what point does self-hosting on vLLM beat a managed LLM API?
Short Answer
Medium
Premium
How does semantic caching…
Short Answer
Hard
What does prefix caching store, and which workloads benefit most?
Short Answer
Medium
A timestamp at the top of every prompt tanks the cache hit rate: find the mistake
Spot the Error
Medium
Is trimming the prompt usually the biggest cost lever on a chat endpoint?
Multiple Choice
Medium
Select the levers that cut output token spend without hurting answer quality
Multi-select
Medium
Premium
Design a routing tier…
Short Answer
Medium
Which signal best decides when to escalate a query to a larger model?
Multiple Choice
Medium
Price this call: 1,500 input tokens and 600 output tokens on a given rate card
Predict Output
Medium
When an async batch API beats the real time endpoint
Flashcard
Easy
Select all dimensions that meaningfully drive cost on a managed vector database bill in 2026.
Multi-select
Easy
When does inflating the system prompt actually lower cost per call?
Short Answer
Medium
'Prompt caching cuts our bill by 90 percent across the board', what's overclaimed?
Spot the Error
Medium
A teammate randomized system prompts per user and expects the prompt cache to still help, what's wrong?
Spot the Error
Medium
Greedy decoding and temperature sampling: does one cost more per token than the other?
Short Answer
Medium
Bytes, characters, tokens: which one gets billed and how do they relate?
Flashcard
Easy
Why does a reasoning model like DeepSeek V4 quietly inflate the output token bill?
Short Answer
Medium
Pick Claude Sonnet 4.6 over Opus 4.7, when is that the right call?
Multiple Choice
Medium
Premium
Pick the workload where…
Multiple Choice
Medium
LoRA vs full FT on 7B + 10k examples: order of magnitude cost comparison
Short Answer
Medium
Premium
How do Anthropic and…
Short Answer
Medium
Which prompt construction pattern is most likely to actually hit the provider's prompt cache?
Multiple Choice
Medium
Showing 1–25 of 28
← Prev
Next →
Evenup
Sourcegraph
Servicenow
Sharechat
Capgemini
Databricks
Anthropic
Linkedin
Anthropic
Mckinsey
Anthropic
OpenAI
Inflection Ai
Infosys
Ironclad
Krutrim
Doordash
Promptlayer
Ey
Salesforce
Cerebras
Ola
Anthropic
Coinbase
Alibaba
Character Ai
Goldman Sachs
H2o Ai
Ey
NVIDIA
Bytedance
Stripe
LlamaIndex
Notion
Arize Ai
Scale Ai
Deepseek
Gnani
Dataiku
Elevenlabs
Browserbase
Cognizant
Jpmorgan
Lakera
Pinecone
Tencent
Accenture
Contextual Ai
Comet Ml
Haptik