Skip to content
TrackPodcasts
scienceOct 2, 202622:22

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Get every episode summarized

Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“Today's paper is selected from the Hugging Face Daily Paper List of October 2, 2026, and has received 32 upvotes. The title of the paper is Active Sabbler, Automated Curriculum Learning for Agent Harness Optimization.”From the transcript

🤗 Upvotes: 32 | cs.AI, cs.CL, cs.LG, cs.MA, cs.SE

Authors:
Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor Rühle

Title:
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Arxiv:
http://arxiv.org/abs/2610.00906v1

Abstract:
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.

Know when Sungho Park turns up

Follow Sungho Park and once a week we email you every new episode they appeared on — including guest spots the show notes never mention, because we read the transcript.

Follow Sungho Park

Free. Pick your own day and time.

Hosts & guests

Transcript ready

463 searchable segments. Every word is indexed and playable.

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Daily Paper Cast

0:00
22:22

Full transcript

Daily Paper Cast — ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Welcome to Daily Papercast. Today's paper is selected from the Hugging Face Daily Paper List of October 2, 2026, and has received 32 upvotes. The title of the paper is Active Sabbler, Automated Curriculum Learning for Agent Harness Optimization. The first two authors of this paper are Sun Go Park and Juan Joong Kim, with the corresponding author being Zhu Aijong from Microsoft. All right, let's dive into the introduction of this research. Large language models, or LLMs, have become increasingly capable of powering autonomous agents, enabling them to handle complex tasks. However, achieving reliable performance, especially on multi-step and long horizon tasks, is still highly dependent on the external harness surrounding the model. External harness? What exactly does that entail? A harness in this context specifies how the agent is prompted,

which tools it can use, and how its execution is controlled, monitored, and corrected. These aspects are crucial for improving the robustness and the overall effectiveness of the agents. I see. So a robust harness is vital for the performance of these agents. But what challenges are addressed in this paper? The major challenge addressed in this paper is that existing methods typically focus on how to update the harness from feedback obtained through execution traces. However, they mostly use predetermined sets of training scenarios to generate this feedback, which doesn't adapt to the evolving state of the harness itself. But why is it important for the training scenarios to adapt alongside the harness evolution? As the harness evolves, the most useful scenarios for further optimization can change. Unresolved failures might benefit from further optimization, while already repaired ones might need less attention. A fixed scenario order can't adapt to these shifts, potentially leading to inefficiencies

where unresolved failures remain unfixed, or roles are spent on scenarios that no longer provide useful feedback. That's interesting. So a fixed schedule might either move on to quickly from unresolved issues, or spend too much time on already addressed ones. How does this paper propose to solve that problem? This paper introduces ActiveSadler, which models this scenario selection as an automated curriculum learning problem. ActiveSadler treats the evolving curriculum as a non-stationary bandit problem, dynamically targeting specific optimization goals. Non-stationary bandit problem? How does this exactly work in practice? Essentially, ActiveSadler abstracts recurring failures into reusable failure patterns known as arms. It estimates the potential learning progress provided by further targeting each pattern and adaptively balances revisiting known weaknesses with exploring new unseen scenarios for potential issues.

This adaptive optimization continuously updates both the discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. So it's not just about fixing what's broken, but also finding new points of failure systematically and efficiently. Can you tell us about the main contributions of this paper? The paper demonstrates that ActiveSadler consistently discovers stronger harnesses across different benchmarks, specifically Gaia 2 and Terminal Bench 2.0. In terms of metrics, it improves test pass at 1 by 4.4 and 7.5 percentage points, respectively, over the same harness optimizer using a fixed scenario order. Ablation studies conducted further validate that the gains achieved depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing optimization with new failure discovery. It sounds like ActiveSadler represents a major step forward in automated harness optimization, optimizing not just the harness,

but also continuously adapting the training scenarios. That effectively addresses the evolving nature of the tasks and models. Indeed, and these results establish automated curriculum learning as a crucial dimension for harness optimization. The paper argues that performance improvement depends not only on how feedback is used for updates, but also on which training scenarios generate that feedback. Great insights. That's the end of the introduction section. Now we have a solid understanding of the context and objectives of ActiveSadler. Now that we have a good grasp on the introduction and background of ActiveSadler, let's delve into the methodology. Ashley, can you walk us through the methods used in this paper? Evan, the methodology section of this paper begins by framing the problem formulation and the design rationale behind ActiveSadler. Let's dive right into it. All right, ActiveSadler aims to optimize the training scenarios to improve the harness of an LLM agent.

It's designed as an iterative curriculum learning framework that continuously adjusts the scenarios based on performance feedback. So how does it decide which scenario to use next for optimization? Good question. The process is modeled as a sequential resource allocation problem, framed as a multi-armed bandit. Let's break that down a bit. In this context, each arm represents an optimization target, which is essentially a pattern derived from recurring failures. The system then decides which arm to pull or select next for optimization. That makes sense. How do they ensure that the system keeps discovering new weaknesses and doesn't just focus on the existing known issues? Exactly. This balance is crucial. And this is where the exploration controller comes into play. It decides whether the system should explore unseen scenarios, which might expose new failure patterns, or exploit known arms to optimize existing issues. The decision-making is dynamic and based

on the current state of the harness and the optimizer's history. All right, so the system can choose to explore new scenarios or continue to work on known issues. What comes next in the workflow? Once the exploration controller makes its decision, the next component is the arm prioritizer. This module is responsible for scoring the current arms or failure patterns based on their estimated learning progress. It uses a learning progress score denoted as fit to prioritize which arms should be targeted for optimization. Can you give more details about this learning progress score? Sure. The learning progress score is calculated by considering four main factors for each arm. Severity, fixability, breadth, and side-effect risk. Essentially, these factors assess how impactful the failure is, how easy it is to address, how broadly a fix would apply across different scenarios, and the potential risk of the fix causing regressions or other side-effects. Got it.

And after scoring the arms, what does the system do next? Based on these scores, active Sadler uses a stochastic selection mechanism to decide which arm to pull next. This approach prioritizes higher scoring arms while still allowing for the possibility of selecting lower scoring ones if their scores improve over time. Interesting. And once an arm is selected, what happens next? Once an arm is selected, the next batch of scenarios associated with that arm is executed under the current harness. The execution results, which include success or failure feedback, are then fed back into the system to update the learning progress scores and the state of the arm. So it's a continuous feedback loop that helps the system to adapt and evolve. What about the scenarios that get newly discovered? Exactly, Evan. Newly discovered failure patterns are instantiated as new arms. These are then included in the pool and prioritized accordingly in subsequent iterations.

Active Sadler also ensures that if an arm has been resolved, it gradually receives lower priority, preventing over-optimization on already fixed issues. This sounds like an efficient way to ensure the system is always working on the most relevant issues. How does the failure pattern extractor fit into this workflow? Great point. The failure pattern extractor operates at two stages, symptom extraction and pattern normalization. It first abstracts each failure into a candidate symptom without considering the existing arms. After all candidate symptoms are extracted, these are normalized against the current arms, resulting in matched, composed, or new arms being instantiated. So the failure pattern extractor essentially ensures that new issues are modularly identified and incorporated into the process. Precisely. This modular identification is key as it allows recurring issues to be efficiently tracked and addressed.

Quick-ablation studies and evaluations further showcased active-sabler's effectiveness. Can you touch upon the role of these evaluations briefly? Ablation studies demonstrated that active-sabler's performance significantly relies on three principles. I, using failure pattern arms effectively, E, adaptive arm prioritization, and three, the exploration exploitation balance maintained by the exploration controller. It seems like every component works in synergy to ensure optimal performance of the agent harness, making the optimization framework both adaptive and robust. Indeed, this constant adaptation is what allows active-sabler to effectively improve the harness over time, responding both to newly discovered and evolving patterns. That's an excellent rundown of the methodologies used in active-sabler. Comprehensive and intricate, yet logical. Yes, and that wraps up the method section of the paper.

The next segment will cover the experiments and their results. All right, now that we have delved into the methodology behind active-sabler, let's turn our attention to the experiments and results. Ashley, could you guide us through the experimental evaluation conducted by the authors? Certainly, Evan. The authors evaluated active-sabler on two benchmarks, Gaia 2 and Terminal Bench 2.0, which covered diverse domains and task types. Can you give us more context about these benchmarks? Of course. Gaia 2 evaluates assistant capabilities in a simulated smartphone environment across 10 distinct user environments, or universes. Terminal Bench 2.0 comprises 89 realistic tasks in areas such as system administration and cybersecurity. Both benchmarks aim to test the harness optimization in varying settings. How is the experimental setup designed for these benchmarks? For the evaluation, the authors use the same train,

development, and test splits as autosabler. Gaia 2's tasks were partitioned by universe, ensuring that each split contained tasks from distinct user personas, whereas Terminal Bench 2.0 used a uniform random partition, given its varied task domains. That sounds thorough. So how does active-sabler perform compared to other methods? Active-sabler showed the strongest test performance on both benchmarks. On Gaia 2, it achieved a pass at one of 59.8%, compared to 55.4% for autosabler and 54.2% for GIPA. Similarly, on Terminal Bench 2.0, active-sabler reached 80.0%, while autosabler scored 72.5%, and the manual-benchmark scored 69.2%. Impressive improvements. What are the Ablation Studies reveal about the contributions of different components in active-sabler? The Ablation's highlight that the gains

depend significantly on three core components. Failure-pattern arms, adaptive-arm prioritization, and exploration control. When failure-pattern arms were replaced with category or scenario-level arms, the test pass at one dropped on both benchmarks, underscoring their importance. How about the scoring and exploration mechanisms? Removing adaptive-arm scoring and replacing it with fixed prioritization diminished performance by 4.5% points on Gaia 2 and 7.5 on Terminal Bench 2.0. Similarly, replacing adaptive exploration with a fixed schedule reduced test pass at one by 4.5 points on Gaia 2 and 8.3 on Terminal Bench 2.0. These findings affirm that dynamic prioritization and exploration are crucial for optimal performance. So the dynamic elements of active-sabler really drive its success? Were there any other significant results from their experiments?

They also demonstrated that active-sabler achieves a better end-to-end cost-accuracy trade-off. Despite higher optimizer overhead, its efficient allocation of task agent rollouts leads to substantial gains in overall cost-efficiency. Cost-efficiency sounds important. Can you elaborate on how this was measured? Sure. Both optimizer-side computation and task agent execution were profiled. On Gaia 2, active-sabler reached 58.5% development accuracy at $298, whereas autosabler needed $1360 for the same accuracy. On Terminal Bench 2.0, active-sabler achieved $78.9 at $128 compared to autosabler requiring $220. It seems active-sabler not only improves effectiveness, but also ensures practical cost savings, making it viable for real-world applications. The method ensures that expensive rollouts

are allocated to optimization targets with the highest potential, balancing between fixing known issues and discovering new ones. That's a comprehensive evaluation. Active-sabler systematically addresses the evolving needs of harness optimization through adaptive learning and cost-effective measures. Exactly, Evan. And that brings us to the end of the experiment section. The next segment will cover the related work and further discussions. All right. With a solid understanding of how the experiments were conducted and their impressive results, let's shift our focus to the related work that lays the foundation for the advancements made by active-sabler. Ashley, what does the papers say about the related work in this area? Certainly, Evan. The related work section of the paper is quite comprehensive. It covers several key areas, starting with automatic harness optimization or AHO. What exactly is automatic harness optimization?

Automatic harness optimization encompasses recent work that goes beyond just optimizing prompts. It extends to the broader agent harness, including tools, memory, skills, and runtime control logic. The general idea is to use rollout feedback to propose updates and evaluate candidate harnesses. Interesting. How do these methods typically operate? These methods can be grouped into search-based and learning-based approaches. Search-based methods explore alternative harness implementations or configurations through techniques like evolutionary or Bayesian search. Meanwhile, learning-based methods treat the external harness state as the object being learned, updating it repeatedly from rollout feedback on training mini-batches. I see. So there's a broad spectrum of methods at play here. Exactly. For example, GIPA is a hybrid of these styles, coupling reflective updates from training mini-batches with evolutionary Pareto search over validation performance.

And how do learning-based methods stand out in this context? Learning-based methods are akin to model training, but without updating the model weights. They repeatedly adjust the harness based on training feedback, aiming to improve the external components around the model. Does the paper distinguish between online and offline optimization approaches? Yes, it does. Offline methods optimize the harness before deployment using training data and held out validation, while online approaches continue adapting during sequential interaction or evaluation. That's quite a landscape of techniques. What about automated curriculum learning? Automated curriculum learning is another crucial area discussed. It studies how training data, tasks, or environments are scheduled to improve learning outcomes. This concept prioritizes learning opportunities according to competence or learning progress. Can you give us some concrete examples of automated curriculum learning?

Certainly. Early work in this area includes algorithms that used non-stationary bandits. Teacher student tasks selection and progress-driven environment sampling. These methods estimate task utility from observable learner side signals such as loss reduction or competence progress. So, active-sabler builds on these principles, but applies them to harness optimization. Exactly. It adapts this principle by prioritizing training opportunities according to their evolving value to the learner. In this case, the harness. It constructs failure pattern arms from diagnosed weaknesses and continually reassesses their optimization value. How about task allocation for agent optimization? Are there adaptive methods for that as well? Definitely. Adaptive task selection has been applied in different stages of agent optimization. For instance, co-evolve uses forgetting and uncertainty observed in agent rollouts to generate new training tasks targeting behaviors the current agent struggles with.

Are there techniques specifically for harness optimization? Yes. Task co-evolve adapts validation tasks to best distinguish candidate harnesses, while harness lens selects verification tasks relevant to each proposed harness modification. These methods adapt tasks for evaluation, whereas active-sabler adapts the training scenarios generating evidence for updates. So, active-sabler fits within this broader context of adapting optimization strategies to the evolving state of the agent or harness. Precisely. It uniquely combines several elements from these existing methods, automated curriculum learning, adaptive task allocation, and automated harness optimization to dynamically adjust training scenarios and leverage failure patterns effectively. That's a thorough overview of the related work. It's clear how active-sabler stands on the shoulders of extensive research while pushing the envelope in adaptive harness optimization.

Indeed, the paper situates itself well within the fields of automatic harness optimization and automated curriculum learning, making significant strides in both areas. And that brings us to the end of the related work section. Stay tuned as we dive into the discussion and further implications of their findings in the next part. With the related work covered, let's summarize the key contributions and takeaways of the paper. Ashley, could you recap for us? Certainly, Evan. The core contribution of active-sabler is its innovative approach to automated curriculum learning for agent harness optimization. By modeling scenario selection as a non-stationary bandit problem, active-sabler dynamically targets optimization goals and adapts the training curriculum alongside the evolving harness. And what sets active-sabler apart from existing methods? Three main components distinguish active-sabler, failure pattern arms, adaptive arm prioritization,

and dynamic exploration control. Each of these contributes significantly to the system's ability to continually optimize harness performance by balancing between fixing known issues and discovering new ones. That's right. The paper also demonstrated impressive results in terms of metrics. The implementation of active-sabler on benchmarks like Gaia 2 and Terminal Bench 2.0 showed substantial improvements with test pass at 1, scores increasing by 4.4 and 7.5 percentage points, respectively. Moreover, active-sabler achieved better end-to-end cost efficiency, efficiently allocating task agent rollouts to optimization targets with the highest potential. This means not only improved accuracy, but also practical cost savings, making the approach highly viable for real world applications. In essence, active-sabler represents a significant leap in harness optimization. It ensures continuous adaptation and improvement,

addressing the evolving needs of complex LLM agents in varied task environments. The paper firmly establishes automated curriculum learning as an essential dimension for optimizing agent harnesses, highlighting that effective feedback usage and adaptive scenario selection are pivotal for achieving robust performance. And that wraps up today's episode. Thank you for joining us on Daily Papercast. We hope you found this discussion as enlightening as we did. Don't forget to subscribe and tune in daily for more insights into the latest advancements in AI, NLP, and CV research. See you in the next episode. Have a great day, everyone.

More episodes

More from Daily Paper Cast

View all episodes →