Post

Log inSign up

Post

Log inSign up

Fuli Luo on X: "Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: https://t.co/ZSxahzJRju"

@_LuoFuli
Fuli Luo
@_LuoFuli
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/
6:52 PM · Sep 16, 2026·
3.1M
Views
424

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email

Relevant people

Avatar
Fuli Luo@_LuoFuliFollow
Now building @XiaomiMiMo. Previously @deepseek_ai

Trending now

Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @_LuoFuli
    Fuli Luo
    @_LuoFuli
    Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/
    6:52 PM · Sep 16, 2026·
    3.1M
    Views
    424
  • @_LuoFuli
    Fuli Luo
    @_LuoFuli
    Sep 16
    We believe RL is one of the most scalable and efficient paths toward self-improvement.
    12
  • @_LuoFuli
    Fuli Luo
    @_LuoFuli
    Sep 17
    We hope this livestream sparks the research community’s interest in the core challenges of scaling RL and encourages researchers to help us refine our training recipe. We’ve been delighted to see so much insightful analysis based on the detailed training metrics.
    5
  • @hxiao
    Han Xiao
    @hxiao
    Sep 16
    very nice and open! 🫡 just dont let cfo see this
    GIF
    4