Serverless GPU compute platform optimized for LLM inference with fast cold starts (sub-10s) via memory snapshots, supporting vLLM/TGI and per-call billing.
18 more curated Modal questions: mixed formats, calibrated difficulty.