← Home
Blogs
Writing, research, and case studies from Monte.
- Improving Long-Horizon RL with Process-Level Rewards As agentic task horizons grow longer, constructing process-level reward signals becomes a necessity for effective RL training.
- The Case for Smaller Models Frontier model prices keep rising. We argue that fine-tuning small models for specific workflow steps is not only cheaper than calling a frontier model, but for many agents, the only way to get reliable results in production.