-
OptiMind: Teaching LLMs to Think Like Optimization Experts
Paper • 2509.22979 • Published • 4 -
LFM2 Technical Report
Paper • 2511.23404 • Published • 56 -
Zero-Overhead Introspection for Adaptive Test-Time Compute
Paper • 2512.01457 • Published • 2 -
Confidence Estimation for LLMs in Multi-turn Interactions
Paper • 2601.02179 • Published • 17
Collections
Discover the best community collections!
Collections including paper arxiv:2601.06487
-
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
Paper • 2512.24618 • Published • 154 -
Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem
Paper • 2512.24873 • Published • 108 -
AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents
Paper • 2512.23343 • Published • 30 -
Figure It Out: Improving the Frontier of Reasoning with Active Visual Thinking
Paper • 2512.24297 • Published • 6
-
lusxvr/nanoVLM-222M
Image-Text-to-Text • 0.2B • Updated • 208 • 99 -
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Paper • 2503.09516 • Published • 39 -
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
Paper • 2505.24863 • Published • 97 -
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Paper • 2505.17667 • Published • 88
-
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
Paper • 2511.21087 • Published • 10 -
Nested Browser-Use Learning for Agentic Information Seeking
Paper • 2512.23647 • Published • 19 -
KV-Embedding: Training-free Text Embedding via Internal KV Re-routing in Decoder-only LLMs
Paper • 2601.01046 • Published • 14 -
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
Paper • 2601.03193 • Published • 50
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 104 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
OptiMind: Teaching LLMs to Think Like Optimization Experts
Paper • 2509.22979 • Published • 4 -
LFM2 Technical Report
Paper • 2511.23404 • Published • 56 -
Zero-Overhead Introspection for Adaptive Test-Time Compute
Paper • 2512.01457 • Published • 2 -
Confidence Estimation for LLMs in Multi-turn Interactions
Paper • 2601.02179 • Published • 17
-
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
Paper • 2512.24618 • Published • 154 -
Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem
Paper • 2512.24873 • Published • 108 -
AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents
Paper • 2512.23343 • Published • 30 -
Figure It Out: Improving the Frontier of Reasoning with Active Visual Thinking
Paper • 2512.24297 • Published • 6
-
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
Paper • 2511.21087 • Published • 10 -
Nested Browser-Use Learning for Agentic Information Seeking
Paper • 2512.23647 • Published • 19 -
KV-Embedding: Training-free Text Embedding via Internal KV Re-routing in Decoder-only LLMs
Paper • 2601.01046 • Published • 14 -
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
Paper • 2601.03193 • Published • 50
-
lusxvr/nanoVLM-222M
Image-Text-to-Text • 0.2B • Updated • 208 • 99 -
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Paper • 2503.09516 • Published • 39 -
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
Paper • 2505.24863 • Published • 97 -
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Paper • 2505.17667 • Published • 88
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 104 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75