Publications
* Equal contribution † Project lead The leading papers are highlighted.
AI Can Learn Scientific Taste
- Introduces Reinforcement Learning from Community Feedback (RLCF), using 720K field- and time-matched citation pairs to learn community preferences for high-impact research.
- Trains Scientific Judge and Scientific Thinker for impact assessment and ideation; the 30B judge reaches 82.7% accuracy, while the thinker outperforms baselines on idea potential.
Self-Foveate: Enhancing Diversity and Difficulty of Synthesized Instructions from Unsupervised Text via Multi-Level Foveation
- Introduces Self-Foveate, a “Micro-Scatter-Macro” pipeline that extracts fine-grained details, cross-region relations, and holistic patterns from unsupervised text.
- Adds iterative re-synthesis to improve source fidelity; experiments across diverse settings show consistent gains in instruction diversity, difficulty, and downstream performance.
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- Proposes “Thinking with Video” and VideoThinkBench, using generated video frames as a shared medium for dynamic visual and text-centric multimodal reasoning.
- Shows Sora-2 is competitive with leading VLMs on visual tasks and reaches 92.0% on MATH and 69.2% on MMMU; few-shot learning and self-consistency further improve reasoning.
STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
- Proposes STAR-S, an iterative self-taught loop that elicits safety-rule reasoning, repairs failures through guided reflection, and fine-tunes on the resulting traces.
- Across six jailbreak and two over-refusal benchmarks, STAR-S outperforms safety-alignment baselines while balancing over-refusal and preserving general capabilities.



