Tren
dar
Dashboard
Papers
News
GitHub
Updates
KO
EN
Sign in
Updates
Papers
Dashboard
News
GitHub
“reinforcement learning”
Papers, GitHub repos, and news related to this keyword, in one place.
Papers
12
All →
Semantic Scholar
NLP · LLMs
530 citations
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
OpenAlex
NLP · LLMs
1.4K citations
A Survey of Large Language Models
Semantic Scholar
Multimodal
568 citations
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Semantic Scholar
NLP · LLMs
395 citations
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Semantic Scholar
NLP · LLMs
264 citations
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Semantic Scholar
NLP · LLMs
233 citations
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Semantic Scholar
NLP · LLMs
152 citations
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
Semantic Scholar
NLP · LLMs
105 citations
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Semantic Scholar
NLP · LLMs
12 citations
Evaluating large language models for accuracy incentivizes hallucinations
Semantic Scholar
ML Methods
8 citations
Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
Semantic Scholar
ML Methods
8 citations
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
Semantic Scholar
Multimodal
3 citations
Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models
GitHub
1
All →
AI · Finance
Jupyter Notebook
★ 15.7K
AI4Finance-Foundation/FinRL
News
4
All →
Hacker News
Safety
▲ 9
Reinforcement Learning with Metacognitive Feedback Elicits Uncertainty in LLMs
Hacker News
Reinforcement Learning
▲ 6
TycoonLE: A Jax reinforcement learning environment for long-horizon planning
Hacker News
▲ 74
The Little Book of Reinforcement Learning
arstechnica
▲ 0
Quantum error correction can constantly recalibrate a processor