Zenaique

Spot the errors in this R1-Zero vs full R1 comparison.

Spot the error·Medium·4.0 · 0·~2 min·Asked atComet MlTcsUnity·Relevant atGoogleMetaOpenAI
Attempt it

Click any words you think contain an error. Click again to unmark.

Mark at least one word to submit.
TL;DR

Two errors: R1-Zero traces are messy multilingual CoT, not polished English for users; cold-start SFT seeds readable format and coverage — not marketing copy.

Memory aid
Sign in to see the mnemonic that makes this stick.
Easy to grasp

The passage gets the experiment right but flips the homework quality. R1-Zero's notebook is messy scribbles in mixed languages — fine for proving the student can solve the puzzle. Full R1's extra training lesson is neat step-by-step examples so the final write-up is readable. That lesson is not about ad slogans; it is about how the thinking looks on the page.

Concept explanation~2 min read

Everything you need to truly understand this topic: intuition, mechanics, step by step explanation, code, formulas, and worked example. Click to expand.

DeepSeek's two-stage story — R1-Zero then full R1 — is easy to garble. The spot-error passage gets the scientific headline right but botches what R1-Zero outputs look like and why full R1 reintroduces supervised data. Precision on those details separates candidates who read the report from those who read the tweet.

This walkthrough marks each span correct or false and explains the training logic behind the corrections.

The true claim: pure RLVR without SFT

Sentence one accurately states R1-Zero's contribution: large-scale RL with verifiable rewards can induce strong math reasoning without an initial SFT stage. That was the novel experimental result — not incremental benchmark tuning.

Do not mark this sentence as an error in spot-error drills. Many passages mix one correct and multiple false statements; partial credit in real interviews means naming exactly which clauses fail.

The correct opener sets up why the next clauses are surprising if false: if RL alone works, naive readers assume outputs are also production-polished. They are not.

Error 1: trace quality misdescribed
Error 2: cold-start SFT dismissed as marketing
Putting R1-Zero and full R1 on one timeline
Sign in to unlock the full deep dive.

Situations where this technique stops working.

Sign in to see when this approach fails.

2–4 min · Everything important, quickly.

Sign in to see the quick scan of the deep dive.

Real products, models, and research that use this idea.

  • DeepSeek-R1 report shows R1-Zero trace examples with language mixing vs cleaned full R1 outputs.
  • Distilled Qwen-R1 models use full R1 teacher traces, not R1-Zero raw emergent text.
Sign in to see more production examples.

What an interviewer would ask next. Try answering before peeking at the approach.

QWhy would RL not naturally learn readable English if users need it?
A

Verifier reward omits presentation — unless shaped, readability is invisible to GRPO objective.

2 more follow-ups an interviewer would ask next. Sign in to reveal them.

Red flags & common mistakes

The phrases that signal junior thinking. Click to expand.

Most common mistake

Believing R1-Zero emits display-ready English CoT — emergent RL traces are messy; cold-start SFT fixes format, not marketing.

Sign in to see all red flags and common mistakes.

60 second bullets to scan on the way to the call.

  • Confirm first sentence is true

  • Correct error 1 about polished English traces

Sign in to unlock the revision sheet.

Primary sources. Browse if you want the original framing.

Similar questions

Same topic, related formats. Practice these next.

4 curated
Next question
Match each RL algorithm trait to PPO or GRPO.
Match pairs·Medium