Reinforcement Learning Example Code

Nous Research's NousCoder-14B is an open-source coding model landing right in the Claude Code moment

B, an open-source AI coding model trained in four days on Nvidia B200 GPUs, publishing its full reinforcement-learning stack as Claude Code hype underscores the accelerating race to automate software ...

GitHub

CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning

We are excited to release the CapRL 2.0 series: CapRL-Qwen3VL-2B and CapRL-Qwen3VL-4B. These models feature fewer parameters while delivering even more powerful captioning performance. Notably, ...

True agentic AI is years away - here's why and how we get there

Today's AI agents are a primitive approximation of what agents are meant to be. True agentic AI requires serious advances in reinforcement learning and complex memory.

10d

New framework simplifies the complex landscape of agentic AI

A practical guide to the four strategies of agentic adaptation, from "plug-and-play" components to full model retraining.

eLife

A differentiable model for optimizing the genetic drivers of synaptogenesis

This study presents SynaptoGen, a differentiable extension of connectome models that links gene expression, protein-protein interaction probabilities, synaptic multiplicity, and synaptic weights, and ...

IEEE

Aligning Crowd-Sourced Human Feedback for Reinforcement Learning on Code Generation by Large Language Models

Abstract: This paper studies how AI-assisted programming and large language models (LLM) improve software developers' ability via AI tools (LLM agents) like Github Copilot and Amazon CodeWhisperer, ...

marktechpost

This AI Paper from Stanford and Harvard Explains Why Most ‘Agentic AI’ Systems Feel Impressive in Demos and then Completely Fall Apart in Real Use

Agentic AI systems sit on top of large language models and connect to tools, memory, and external environments. They already support scientific discovery, software development, and clinical research, ...

IEEE

Show inaccessible results

Nous Research's NousCoder-14B is an open-source coding model landing right in the Claude Code moment

CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning

True agentic AI is years away - here's why and how we get there

New framework simplifies the complex landscape of agentic AI

A differentiable model for optimizing the genetic drivers of synaptogenesis

Aligning Crowd-Sourced Human Feedback for Reinforcement Learning on Code Generation by Large Language Models

This AI Paper from Stanford and Harvard Explains Why Most ‘Agentic AI’ Systems Feel Impressive in Demos and then Completely Fall Apart in Real Use

Multi-Agent Evolutionary Reinforcement Learning Based on Cooperative Games

Supervised learning example explained with real-life use case

How to Help a Student's Commitment to Learn

Agent Lightning: Adding reinforcement learning to AI agents without code rewrites