
Agentic Loops for Knowledge Workers
About this episode
The AI Daily Brief: Artificial Intelligence News and Analysis is made possible by:
In this episode, NLW and Nufar Gaspar explain how knowledge workers can move beyond one-shot prompting and use agentic loops to produce more complete, reliable work. They break down how to design verifiable finish lines, decide which tasks should be looped, prevent runaway costs and compose multiple agents into work graphs that can research, review and refine outputs autonomously.
NEXT COHORT - Executive Agent Leadership - Returns in September -- Learn how to use agents - https://training.besuper.ai/
Brought to you by:
KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at https://kpmg.com/us/Sophisticated
Harbor - Invest in the AI ecosystem. https://www.harborcapital.com/aidaily
Hyperagent - Hire a team of always-on agents. New users get $100 in free credits. hyperagent.com/aidailybrief
Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack https://www.rackspace.com/
Section - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/
Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/
The AI Daily Brief helps you understand the most important news and discussions in AI.
Interested in sponsoring the show? [email protected]
Get every episode summarized
Each time The AI Daily Brief: Artificial Intelligence News and Analysis publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
581 searchable segments. Every word is indexed and playable.
Full transcript
The AI Daily Brief: Artificial Intelligence News and Analysis — Agentic Loops for Knowledge Workers. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Throughout the summer, one of the hot topics among advanced AI users has been the idea of loops or loop engineering. Simply put, the concept is to think about the way that we interact with AI, not as prompting it and telling it what to do, but to setting up the circumstances where the AI or agent can loop over and over again, working to complete a specific task, with a measurable output that it can check itself against, running until that task is complete based on that measurable goal. The first place loops to cold was of course in software engineering, where the nature of the tasks is fairly definable and success is pretty clear. Moving loops into knowledge work domains, where sometimes success is less definable, is more of a challenge, but it's not impossible. If you have the right tools to design your knowledge work tasks for this type of agentic work. Today's episode is a webinar with Newfar Gaspar where we do exactly that, and that is coming up right now. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
Alright friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Harbor and HyperAgent. To get an ad free version of the show, go to patreon.com such a.i.deally brief or you can subscribe on Apple Podcasts to learn more about sponsoring the show. Send us a note at sponsors at aideallybrief.ai. And if you like Newfar's presentation on this and you want to go deeper into the world of building agents, allow me to recommend our super intelligent executive agent leadership program. The next cohort is kicking off next week. It is led by NewFar and you can find out all about it at training.besuper.ai. Lastly a note, I am traveling currently for Liberty and my birthday. So if something absolutely crazy has happened and you're wondering why the heck you are getting this agentic loops presentation, that is why. Although obviously if there is something big enough, I will pop back in. For now, let's dive into agentic loops for knowledge workers. Today we will cover, I believe some very important topics around loops and graphs and in general how to utilize the most advanced techniques for getting agents to work autonomously and as a group.
And I want to hand it over to Netanyl to set the table stakes as to why we are here today. So one of the really interesting dynamics right now is we're pretty well past the point where people hear you say you're vibe coding or doing something with cloud code and assume that you're now all of a sudden as a knowledge worker trying to become a software engineer. It's very clear that we're kind of in the phase of actually figuring out how code and software engineering style processes and these set of tools can make their way into other aspects of work and influence how that work. It's done. And just a couple of weeks ago, OpenAI dropped these recent usage statistics which show just how dramatically consumption and use of AI has shifted from the assisted to the agentic. So the chart that's on your screen is from that you can see around April, May, we flipped from majority AI usage in terms of total tokens consumed being in that kind of chat GPT assisted paradigm to the
agentic paradigm. The notable thing about this is that alongside this more advanced type of usage, you also see the firms and individuals who are using AI in these new agentic ways are pulling away. They're the space between them and others at least in terms of tokens consumed, you know, which is obviously a pretty rough metric. But at least by that metric, they are getting farther apart from the average. What's really difficult, and I'm sure a lot of you have felt, is that knowing how to translate these concepts that originate in software engineering to other types of knowledge work can be a difficult process. It can be abstract. It can involve layers of abstraction. And so we wanted to put together this webinar because this idea of using agents in loops and sort of no longer prompting but designing loops and things like that has been kind of part of the buzzy zeitgeist of AI early adopters for a few months now. But I think it's still been remained abstract what it actually means for knowledge workers. So that's the goal of this is to kind of bring everyone into
new ways of interacting with AI. And I'm excited to see where we go with it. All right. So Alup is basically a job and a graphic and an organization. That's a quote actually from you from the podcast. Keep that in mind and we'll walk you through that. But if I need to give you like a TLDR of what we're doing in three sentences. So the first thing is that AI tools, the ones that most of you are using the agent tools already use loops to do the work behind the scene. And once you learn to give them a concrete and verifiable end goal, they will keep working until the job is actually done without you nudging them or without you being frustrated by the mediocre result potentially. That's the big promise. Okay. And when one loop and one agent stops being enough, you can always compose loops into teams of agents. And that's the whole idea behind the Graph Engineering noise and the chatter in social. There is substance around that. And this is very important. All of these ideas were born in software engineering. So if you are coming from this background or
have computer centers working for you and your company and so on, the graphs are not news to any of them. And the engineers have been building with these concepts for a couple of years and over the last few months with a lot of focus on specifically on agents. But almost all the practice as Netanyl said is very focused on coding. And we will try to show you the pros and cons, the pitfalls and how best to leverage that to other types of work, namely knowledge work. One last thing to say because most of the work that you do regularly probably requires just let's call it regular agent execution. And only in some cases you might need loops, the graph and the more sophisticated things should be for special occasions. Meaning I want you to keep things simple and go into loops and graphs and multiple agent orchestrations only when needed. So I want you to keep it simple while being mindful of the full breadth of how far you can take the modern technology. Very quickly, what happened on social for those who were not sitting in the echo chamber of
Twitter X all the time in July six words by the founder of OpenClaw. What he was basically saying is that turn no longer talking about loops. We're talking about graphs very quickly, millions of views. He was half joking, but very quickly and obituary for the loop engineering. That is primarily an aiming event, but I think the joke stuck because it pointed into something that was more real and to be very concrete of what's kind of happening or the naming letter that does have some substance around it. This is how we got here. And this is I think what tells the story of the evolution of agents and AI over the last month or even one or two years. So we were very obsessed at the beginning about what you say to the model that was born engineering and we all learned how to speak effectively to the models. Then we got obsessed about what the model knows. That's the context engineering that we talked about extensively. Then we started talking about where it runs and what
tools it can touch and that became the conversation around harness engineering. And this year we started talking more and more about how long the models and the agents can run on their own. That's the loop engineering. And lastly, we started asking how many of them work together in order to get the job done and that's the essence of what is now referred to as graph engineering. If you notice kind of the direction of the travel, each evolution is about given the AI more independence at a bigger scale. That's where we're headed. And some people even claim that this is just a rebranding of a very natural evolution alongside the technologies abilities. And because the field and the practitioner is adding new terms, new skills, evolution, roughly every couple of months and chasing them can be exhaustive. I think the more important thing is the essence is knowing how to get the agents to work effectively and how to orchestrate them. That's the skill that probably survives the every renaming and every buzz on social. At the end of the day, we just want to get the job done and to be ambitious
about the type of jobs that we can do. All right. So that's the background. One last data point that I think proves that loops are important just three weeks ago when Jeff Dean, one of the most renowned engineers I think in the history of modern technology has left Google. He basically went to establish a company literally called Discovery Loop. And this company is going to be around using loops for scientific discoveries. So there is merit about being able to run AI in loops in order to get progressively better results. Okay. So that was the background. Now I want to make sure that we are all on the same page with regards to what is a loop and work and you find one. And the one thing that you need to understand that every agentic tool that you use, whether it's co-work, jet-gpt work, codex, cursor, any agentic tool for that manner, what is being implemented under the hood is already a loop. That's what the harness is doing and in various degrees of effectiveness.
Basically it runs through an iterative process of planning how to get the job done acting, typically using tools, checking whether the results that were received from the tools are sufficient or good enough and then adjusting. The thing is that the loop that was implemented by the companies behind the agentic tools are as good as what they implemented and in general they are quite generic. And that's why even though these tools are amazing and we are able to get very good results out of them, in many cases we are observing the fact that we need to kind of nudge them or we need to be very proactive in our prompting if we want them to work extra hard or run multiple iterations on something before it stops. That's not something that native loop that were already implemented by the companies will not do that for you. But the one thing to understand is that under the hood there is always a loop and then the question what is all this buzzword or what is the loop that advanced loop that we're talking about. So today what we're talking is about a loop that basically you control what is the end goal. So we're looking for a loop that happens when you extend the native cycle,
the cycle that is already run by the tool and the way to extend it is by leveraging the specific commands in most tools it's called a slash goal command in cursor it's called literally loop command and the entire purpose of this is to give the tool a concrete end goal and make sure that this end goal is highly verifiable and making sure that there is a way for the tool to progressively test itself versus is it done or not. Okay so that's the entire thing that we're talking about today and if you want to master two things that really matter here is first of all you need to understand when you need to create this special loop the one that you control the end goal and the one that you encourage the tool to run in multiple cycles until result is done and the second thing that you need to know is to define the correct finish line. That's the most important skill for an knowledge worker that wants to leverage that for their day to day. Two additional notes here. One
is that a loop is not a synonyms for a schedule. A schedule answer the question of when something should run by when it can be based on the clock it can be when something happens for example when there is an email that is being sent that's an automation or a schedule a loop answer the question of until and it stops when the work meets the bar however long that takes or the other constraints that you will give it I'll show those constraints in a minute. So those are profoundly different promises. So don't confuse loops with automations those have different purposes and also the other caveat as I mentioned before because loops were born in coding they are not as optimized for knowledge work as such because coders invented them for coders and for coding and coding has a very clear super power that not all of our work as knowledge worker has it has verification as a very abundant thing that they can execute the code either compiles or it's not compiled the test pass or the
fail it's very relatively easy if there are coders here on the line I don't want to say that your work is easy but in order to verify coding that we have built-in mechanisms which makes the ability to run until a certain condition is being met much more doable okay so when we're talking about knowledge work in many cases we don't have a built-in referee is this report good enough to present to management is this analysis deep enough nobody's compiler answered these questions for you so when someone tells you just put it on a loop they're forgetting that they as coders had free verification and we or you don't have so the good news and that's the central move of what we're doing here today is that verification for knowledge work can be designed by you it's slightly more difficult than for coders but you can manufacture the referee the one that decides whether the job was done and that's what you need to be able to do well and if you are unable to design a
finish line that is very clear and very viable the answer is don't loop it okay so let me give you concrete criteria as to whether the task attend deserves and can be looked for you so in order for a task to be loop worthy or loop relevant it needs to be long-running meaning that it's not something that you can achieve with one prompt and get good enough results with modern models often that's more than enough I urge you to use one shot and getting the results if you can the second thing and those have to come together you have to be able to check whether the results are good enough or the progress is headed in the right direction that's the most important pair the second thing is that you look for things that you probably and potentially want to send overnight meaning that you want to be able to run autonomously maybe over lunch you don't have to run overnight and come back to finish the results instead of a draft that's part of something that might indicate that the loop is
also we want something that should keep running until a specific bar is met or keep running indefinitely watching for something those might also be good indication for a loop we also want something that comes by the way from hard experience is we want something that you tried with the outputs with the fabel with the GPT-SOL whatever smartest model that you have out there and it was just not good enough it didn't meet the bar with the one shot execution also it's relevant for cases where you want to push the model to work harder than one polite pass which is often will be the default think about when you send the model to do the research and that's going to be the example that we will show in a minute often it will run a decent web search synthesize the results and that's going to be that unless you prompt it very vigorously in many cases it will not run multiplications of trying to improve the quality and the abundance of results unless you ask for it very nicely or not so nicely and lastly we want something that has a natural retrain improve shape
okay something that can be a draft but then it gets better with multiplications that's by the way part of I think Jeff Dean is going to do that for research and science because this is a place with more and more and more experiments typically you eventually get to the right direction on the flip side and that's also very important if the task is short and one pass does it if your judgment is the actual work and you cannot upload the judgment to a referee a normal conversation with your agent is the right call and choosing that is the smart move not the copout okay last thing to say loops are the most among the most token hungry execution that you can have with your agentic tools so be mindful that you're using your tokens for the things that matter and that you don't loop for everything we want to look for the things that the value is them all right so here are some concrete use cases of type of knowledge work that people do in order to leverage loops I'll start with the my first attempt of using a loop I ironically decided to use a loop in order to
do a very extensive research on token efficiency best practices that's also something that I will demo in a minute I know that there is a little bit of irony to use the most token wasteful method to look for token efficiency but that's a very good use case that you can pursue research where you want the agent to go deep and wide and synthesize and make sure that you get the results that you're after another very powerful example that many people have been leveraging loops for to do the adding and campaign optimization of highly very viable use case because you can always test with very concrete data whether the click through rate and other analytics that you're using on digital campaigns are actually improving and thereby by iteratively trying multiple things and getting the agents to work on a loop or sometimes indefinitely sometimes with some kind of a cap on how many times it's trying you can overly improve the results and you can see some other results their competitive analysis or scans some audit around content verifying some results doing their
compliance and so on and so forth the places where we're not doing loops is anything that requires the human judgment and cannot be fully automated so if the executive communication requires your judgment it's not a loop similarly with these other folks hiring and strategy because autonomy at the end of the day does not have a taste on its own you probably need to be there to be the final say of that okay so that's a bunch of examples for what people actually looping and if I need to kind of give you the bottom line of the three requirements in order to build a loop so one checkable finish line knowledge work on a six-season loop exactly when you invent a very boring very checkable finish line and boring here by the way is a compliment so for example saying I want 200 verified data points is very boring but very concrete and very machine checkable for you every competitor covered every claim cited summary under 150 words very boring very checkable and you can think about the equivalent
in your domain and you will actually be encouraged to do that in the lab part something like make it insightful is not checkable okay there is no way for the agent to converge on make it insightful okay we're looking for things that the agent can measure itself the second thing that we want to do is to as much as possible have a bounded sandbox it's a space where the loop mistakes are cheap so for example going back maybe to the ad campaigns or to the digital campaigns optimization don't do that on the highest table stakes and let the agent go wild and unless you're okay with the results but for example doing a research in a sandbox or running a specific experiment in a sandbox where if there are mistakes they are not very costly because we're staying in draft mode or in a bounded place that's a better place for you to run the loop and lastly we want a task that can converge so we want use cases where we have different paths go to different potential gates we are in the fact of getting closer to be done research is a task that can converge more
saucers fewer gaps make it better going back to that forever doesn't converge and the agent can get into the loop indefinitely it can always ask itself is it good enough I don't know let me try again is it good enough I don't know let me try again okay so if a task is not well defined and cannot converge we're not going to loop it okay so how do you define that's like the the bottom line how do you define the goal for the loop conceptually you should think about it designing a goal card and these are the things that you should configure as part of the goal cards those also will be the things that you will append after the execution of the actual loop command so it starts with a concrete and clear objective okay the objective should be very clear very machine readable for example in the research that I'm going to trigger in a minute I want to create a definitive token efficiency playbook as of today as of August 2026 you should define what is the output of the loop in our case
it's going to be a file but maybe for you it's going to be something different and this is what makes or breaks everything here because this is how you define the judging and the stopping criteria the initial ones in this case it's going to be I want more than 200 unique data points I want each with URL and data and type I want a specific mix at least 40 vendor docs 40 practitioners 20 benchmarks and I want zero duplications so that's me trying to give the machine a very concrete criteria about what the purpose of the loop and once all of these conditions are met the agent will stop doing the work so your skill is knowing to configure exactly that and if you struggle with defining what's the done when meaning that you're unable to find something that will be very clear very non ambiguous you're going to struggle and you might have a loop running either to show or too long and that's not the right place to go you can it's optional but you can configure stages like specifically what
gates or what type of actions you want the agent to take it's not mandatory sometimes you want to manage that sometimes you actually don't want to do that because you want to leave the agent's efficient amount of judgment to decide how to go about achieving this objective with this stopping criteria and in all of the tools you also have the ability to add an additional let's call them a fail safe or far backs things to avoid the loops running indefinitely so for example here in order to make sure that maybe there aren't 200 unique data points in the internet and the agent will try to run it forever and ever and ever for me unless I give it another stopping criteria so I can tell it try up to 30 turns and sandbox only meaning don't go and do stuff outside the world you can also copy it with time and pair the specific tool that you're using sometimes there are additional things that you can constrain in order to make sure that in case this is something that ends up not being convergent you have a different mechanism to convergent that's actually very important
to have this fail safe mechanism so that's the proper goal let's very quickly try to demo that so again going back to my use case I want to do token efficiency research I'm within as you can see Claude code in Claude code the loop is called slash goal and take a look at my card here research done when all of the following are true the artifact I want the file listed here I want executive summary the data and so on I'm giving it the concrete criteria and giving it a quota it contains at least 200 unique data points about token efficiency and agent decay I work no two points stating the same fact the receipt every data point yeah the adac create a URL this is the mix that I'm asking for as noted and I'm asking for a log I'm asking for a log in order to show you how the look work but I actually think that's a very good practice for you when you're running a loop to ask the model to be a little bit verbose and saying out loud what it's doing and what's the cycle number so you can
see the progression if you want to monitor especially as you're new to using loops it's going to be a great practice for you I'm not going to use favor I'm going to use opus and I'm going to move to the auto mode to make sure that it's actually running now obviously we're not going to sit here and wait for the you know good 20 30 minutes or more that it will take it to run but as you can see go set okay and then it's reflecting the interpretation it will start reflecting the cycles in order to not have you waiting I already ran this exact command earlier and that's what you can see as there is so let's see how it went exact same command and as you can see here so cycle one checked workspace doesn't matter cycle two it got 56 out of 200 it checked various sources in cycle three it got to 90 out of 200 and so on interestingly it went above the 200 in some other executions it went even all
the way to 300 so it's not always as disciplined but fortunately for us we do have the cup of 30 cycles so it will not go above that and that's the like the the finish line and it created an artifact and it's also giving me some things to pull in front of you guys if you're interested so that's how you run a loop here and the other one is running in the background we can check it out later on but that's the entire purpose here so those were loops now I want to make sure that you understand that loops can fail and can fail very miserably the first one is runway spent it just keeps going that's why we have the hard cap and that's why I said it's very critical I've seen people running loops indefinitely or much longer than what they expected the loop to run for we also have a loop that can stack cycles without progress it's sometimes just because it cannot converge or something about the conditions are not there if that's something that happens we need to either stop the execution manually the also dedicated commands in different tools or we can ask it to stop and report
if we know that this is a type of loop because you tried it that sometimes gets stuck sometimes the loop is done but it's very mediocre that's very sneaky because it might have met the letter of your finish line but the results are still very bland so what you need to do here is an acknowledge first of all that it's not the loop failure that's probably the failure of your definition of the goals if it met verbatim the goals that you define but you're not happy with the result that means that you need to better define what is the finish line or what are the referee guidelines and in many cases it's going to be around quality and taste and things that are going to be harder to configure but that's something to note and lastly sometimes we just started running something on a loop that was not something that was meant to be a loop so if that's the case the the turn cap is your best bet. A new study from KPMG in the University of Texas at Austin found that when people work with AI similar skills don't guarantee similar outcomes researchers studied more than 500 early career
professionals and found that the best performers consistently amplify the value of AI by guiding evaluating and refining its outputs these top performers called AI amplifiers weren't defined by what they knew alone but by how they worked with AI learn more about what separates AI amplifiers from everyone else at KPMG dot com slash us slash AI amplifiers. Blitz is deep code based understanding unlocks the thing every road map owner cares about shipping new features here's the truth about building inside a massive enterprise code base writing code was never the bottleneck context is which system does this touch which contracts can't break which standards apply Blitz the already knows because it reverse engineered your entire code base into a dynamic knowledge graph before feature work began with that complete picture Blitz he builds features end to end architecture API's UI and tests all validated against your existing systems one Blitz he customer built an AI native application from scratch with a hundred percent autonomous completion saving over 2700 engineering hours features that respect your code base instead of fighting it stop letting
your backlog grow faster than your team accelerate your roadmap at Blitzy dot com that's BLI TZY dot com every episode I talk about the competition between open AI and tropics SpaceX AI Google and meta and if you've been listening for a while you might have a favorite maybe you think open AI and then the topic can stay ahead or perhaps meta's open source strategy can win out whatever your view every AI lab creates a different investment opportunity harbor capital advisors AI lab ecosystem ETF suite lets you invest in the ecosystem behind the AI lab you believe in search harbor AI lab ecosystem ETFs wherever you invest or follow at harbor capital on X to learn more visit harbor capital dot com for a perspective containing investment objectives risks fees expenses and other important information reading considerate carefully before investing risks include principal loss and artificial intelligence related risks harbor ETFs are distributed by four side fund services LLC harbor is not affiliated with AI daily brief and the funds are not affiliated with sponsored by or endorsed by any AI lab this is a paid advertisement and not personalized investment advice investing involves risk including possible loss of principle this episode of the AI daily
brief is brought to you by hyper agent where you run fleets of agents your team can manage together forget local agents and chat workflows waiting on your laptop to be prompted hyper agent deploys always on agents in the cloud doing real work across the tools your team already uses marketing agents turn competitor moves into landing pages sales agents in rich leads draft emails and updates the CRM ops agent chases the paperwork and tracks the budget every agent has access to shared context and follows your rules about scope and approvals it's time you add agents that feel like teammates hire yours at hyper agent get a hundred dollars in credits at hyper agent dot com slash AI daily brief so we're done with the loop part of the webinar just to summarize we talked about the fact that loops are already a cycle that is being executed by any agent account as that you're using but the loop command or the slash goal command the entire purpose of it is to make sure that you can get the tool to work harder for longer time using a very viable goal also distinguish between
schedule and a loop it was born in coding so you you and we and all of us need to work harder in order to create a finish line that makes sense in knowledge work or not use a loop if we can the main skill here both of these things identify the use case and configure the goal card and always add the caps and the safe sandbox in order to make it work now i want to move from loops to org chart because everything so far was basically one worker it was one agent working alone until done in loops and the progression of this whole field is here okay we all started on the left or if you're you haven't started building agents yet you should be there already then we talked now about looping things when relevant and when doing so well the next station is where the work itself splits basically we have several agents each is doing one piece passing work between them passing information between them that's gonna be a wallcraft and it's built for one job okay and the last
station is when those agents stop being disposable and become your standing team or what is sometimes referred to as an all-graph so one agent working one time agent working on a loop we decompose the work into multiple agents only when it makes sense and when there is a tool need for that more on this to come and lastly if we realize the type of work that we want to get done is much more than the one thing we can build an entire organization of agents as an all-graph to get it to work so maybe one thing to say it's not like that you have to get to number four okay you don't graduate there around many tasks that are more than good for number one and so on but there are some cases where the like the quality of the results and the scale will only be unlocked if you go all the way to building entire teams of agents and orchestrating them efficiently so we only move when there is a justification but in some cases there is huge value to begin here okay so to keep it very simple and concrete I want to explain what a graph is and a graph is basically dots and arrows
that's that okay the dots are called nodes for us a node is an agent or a task that needs to be done the arrows are called edges and the edges include work or information that is flowing and when an arrow has a direction for example research flows into the writing agent we say that the graph is directed and one detail that is very fun and also for the computer science curious or passionate focus on the line look at the bottom line right if we have one node with an arrow pointing back to itself that is the textbook definition of a loop so a loop in a graph are not different things the loop is probably the simplest form of a graph or the smallest form of a graph out okay so the real question isn't always was how many nodes does your work deserve computer science has done work this way for 50 years so it's not new but what is new for our day and age is that AI made the drawing operational today you can draw it and it actually can run for
you okay so that's the big thing here right it's not just something theoretical that you learn in computer science 101 it's something that you can actually execute another the confusion because the same week this stranded half of LinkedIn was also an X of course was also talking about knowledge graph and I've watched very smart people blend three unrelated things into one word so we have knowledge graph those are graph that stores facts what we know and how it connects it's a very beautiful technology completely different jobs not very related to what we're talking here we also have land graph which you may have heard and generous mention this is a developer framework this is a tool for building agent systems in code and it is leveraging the power of graph and a very powerful and good technology and today we're focusing on the work graph it's this is about the execution of the work specifically for knowledge work we're focusing here who does what in what order and what flows between them okay so these are the concepts don't confuse them as much as possible I'm
leaving the computer science engineering alone and I'm talking about graphs and work before everything so let's talk about what we had before AI and before agents the main graph that we actually had and have in each and every organizations concern about the org chart and the org chart the main problem with it is that it basically pretends that work flows in one direction top to bottom right the manager says something the CEO and then it goes to the rest of the organization and that's how we pretended work gets done not at all we all know that work branches it looks back it ends upside ways it ends up sideways it's keeps levels and occasionally flows straight up at 11 p.m. so especially before bold meeting so that's not how work gets done the org chart is the diagram of the authority it's not and never was a diagram of how work gets done graph however dot and L is going wherever the work actually goes is the more relevant picture and the picture that we need to paint to our agents in order for them to follow the work as mentioned to flavor the org graph that
is more persistent and the work graph that is per request or per task okay so let's talk about what changed okay because the skeptics will tell you one thing engineers have been wiring agents in two graphs four years we talked about land graph in a minute ago so why are we also excited about graphs again beyond the the chatter on social the thing that changes the node because the node used to be one fragile lm call a i call and in many cases it was not that great and to orchestrate an entire graph like that was not very feasible for complex task or we have to work very hard in order to get something to work today a node is a whole agent a worker that you hand the job to and the worker has two gears it can be a quick one pass task like we normally ought to typically do or it can be a full loop running until done so either one can be a node and loops are just your heavy duty nodes if you want to kind of understand how everything falls together and what's to all in you here is that
agents got reliable enough to be building blocks within these graphs and that's a big unlock so nothing was invented that is new this summer it's just that something was more democratized and the technology has gone far enough to be able to actually draw this graph on a whiteboard and get it to work effectively and one last very important thing we only go into the orchestration or composing these graphs of multiple agents and multiple nodes when the one worker the one agent that you built is no longer doing the job you're unhappy with the results that you're getting okay so how do you know that a specific task cannot be executed well with a single agent and requires multiple agents and graphs in order to orchestrate in some cases both are valid options if I'm going back to my example of research in many cases a good research agent with a without a loop is more than enough you've seen that after it took about eight cycles right we got a very interesting report out
so in some cases that's good enough if I want to be more comprehensive or I'm discovering that for my intent and purposes it's going to be better to separate the work I can do the same research as a graph meaning as multiple agents that needs to be orchestrated for example I can send multiple parallel agents to research different angles that I'm interested so one can research vendor one can research what practitioners are saying and one can research benchmarks and because these are tasks that are independent I can run them very easily by multiple agents in parallel then I can offload the results into a synthesizer move it to a citation verifier that hopefully is not encumbered by everything that was done prior such that it's an objective verifier and either the human can verify or maybe I would want to learn another verifier by an agent before I publish and I can also maybe add and I will do that an agent that designs the output in a way that I like so
when it is better first of all as I said at the beginning I want you to keep it simple and if one agent works well that's good I want you to move into that in several scenarios first of all scenario which I call a rubber stamp your agent or your loop says done everything checks out and you keep finding issues that it should have caught in many cases self-review is not the way to go especially with things that are critical there is also a lot of work showing that models tend to agree with themselves if you use the same model to verify the results of the same like a GPT verifying GPT odds are it will say that it's correct versus clawed verifying GPT so in many cases we want to fan out to a graph of multiple workers when we realize that the verification is not very reliable another scenario will be that we identify context overflow like one agent is wearing too many hats and it starts confusing them like you are both the objective researcher but also the very creative designer so the researcher might start bringing creative results because it's
time to be creative too early on so when you identify that's the scenario it's probably better to fan out to multiple nodes or multiple agents another thing is where you don't want to sit down and wait for the work to be done seriously and when it can be spawned out to multiple agents during the work parallelly I know that we are all very patient in a this day and age and we're willing to wait for many minutes if not hours to get the results but if we can paralyze the work often we should and another signal will be when the finish line keeps changing mid-run so you keep rewriting basically the goal card while it works because it's really two jobs wearing one card so if you realize that basically it's either an if else or if then kind of a goal that under the hood hides two different goals and two different jobs to be done this is where you probably want to separate and lastly when despite your best effort no matter how much you tried the quality flatline too soon
maybe it's because the context is missing or maybe because it's just not the right architecture you probably want to bring the second perspective or spend out at least one more agent okay if none of this is correct stay in the loop or stay in the agent and don't over complicate thing okay so up until now it sounds very promising like we will build multiple agents they will work in a graph we will describe everything that we have in mind and it will work but the bottom line is how do you actually build a graph what do you need to do and I think that it sometimes sounds over complicated if you talk to the practitioners but in fact there is a progressive level of complexity and they will all work it's just a matter of different competencies and different needs that defines how to do it so the very first thing that you should always do I think but can be a good start is to just do it like a paper a whiteboard you just sketch the text this is very critical because
you need to be clear about how to get the job done and painting that on a whiteboard gets you to confront all of the things that are not well defined and in many organizations many things are not well defined and until you can agree upon a work graph for something you cannot automate that so that's the very first thing that you can do and if you can do it that means that you can even just take a picture of this drawing show it to any agentic tool and it can build it for you so that's why I'm treating that as not only a gateway but as a concrete modality another thing that you already do without knowing or maybe you are paying attention to that but the agentic tools improvise little work graphs every time that you give it a task especially complex task because that's how the harnesses the modern harnesses work you've probably seen some tool maybe it's a cursor or something like that saying I'm going to spawn several sub agents to do different types of work or let me divide and conquer the work and stuff like that so this is under the hood the tool
creating a work graph for you based on the prompt so you're not in control of that but that's another way that the job is doing that and also the co-work and that you should do pretty work they do that as well another way for you to do that is to prompt the graph basically tell the tool research these five competitors in parallel with separate sub agents then have a fresh context reviewer check the merged results against this rule brick that's just one sentence but under the hood but you're describing is a graph that fence out five sub agents to do different research and one to verify that you just build in words seven node graph congratulations another thing that you can do is you can create persistent workers you can create basically either sub agents files or you can create what I refer to as folder agents I'm not going to go deep to all of that because there is a ton of that but I'll show some examples you can basically create these agents as persistent workers that you can summon to the conversation as needed on top of that you can create a skill if you want to have that as a reusable thing that you do or you can just add Hoc
have the tool refer to the specific sub agents that you created we also have the option to create convices using the tools like n8n and other automation tools that lets you basically create a visual graph and lastly we can always create those in code so there are multiple tools and multiple packages that lets you create these graphs and this orchestration via code so this is the five tiers again it's not a competition as to who can get to the highest tier it's just a enumeration of all the ways that you can build these wall graphs using multiple agents doing the work for you for most of the knowledge work that we need to do you will probably live in these tiers and it's more than enough to get the job done let's show you some concrete examples so first of all let's go back to the execution that we had before one thing that I did once the result came in is I said to cloud the following and the playbook to the citation verifier sub agent I'll show you in a minute make sure it's fresh context it knows nothing about how the document was produced and the
document will be that it checks everything there sample 10 random data points and verify each agent against its really URL they then count the totals and the source mix and I'm giving it like a bar so that's me basically taking the result of the loop and getting cloud to add another node and that is the node of citation verification manually just by prompting it and you can see cloud is being very compliant it heading it off-code the verifier gets the file path and the bar nothing about how documents were built to avoid contamination and the verifier verdict is shipped with one required fix and now applied and it gives me the sample verification and the total so you can see how my research went into another node and then I added one more node I want to make it into something visually attractive so I had it added to the report visual sub agent build a beautiful self-contained HTML page from output yada yada it's visual language and attribution rules leaving its role file save
as a output and so on and then it was done so that's one way for you to add additional workers into an existing loop or an existing thing maybe the one of the simplest ways to do that I can also paint a picture on a whiteboard the research sweep that's how I would imagine such a research to happen from here on after and as I said I can just give this picture to cloud and cloud will be able to implement that quite well but the fact that I am able to paint this picture means that I can automate that using multiple agents so that's a very important thing that I can do another thing that I can do I can build a persistent workers okay so what I'm showing you now are sub agents that I created specifically for research I have a benchmark collector that's an agent that brings specific research on benchmark I have citation verifier citation verifier is an agent that the entire purpose is to verify the citation that's what's the one that you've seen in cloud and I have the report visualizer that describes exactly how I want to get things visualized and so on as you've
seen I can summon them to the conversation ad hoc or I can create a skip and the skill here basically describes the graph take a look in case I want to run the research time and again the same way that's a skill research with the work graph as a file and I'm showing it how to do the flow and I'm giving it phases and I'm telling it in each phase of the skill which workers to summon into the conversation and by the way as part of the I don't know if you've seen it but as part of the agents I can also configure which model to use so that's a way that anyone here on the line that knows how to work with cloud knows how to build sub agents cards like you've seen and knows how to build a skill can orchestrate multiple agents very effectively with a lot of controls and it's working with any agentic tool so that's level three basically two more things that I wanted to show you specifically I chose an item but just to make it very visually clear that's another way for you to if you want to create a work rough okay so in this workflow in an item
triggering that on a schedule I have collector agents for different stuff I have synthesizer agent note that ideally I would probably want to use the synthesizer as a strong model whereas the collectors might not have to be a very strong model I have a citation verifier which definitely as much as possible needs to be a different model I can put all of that in a loop until a certain quality is met in this case I added a human verifier and lastly I have a report visualizer and an image so if you are versed in a tool that is more canvas like tool or that's the way you want to work the nice thing about this is that it's very visual and you can very easily see how work flows across the graph and you can verify that so an overkill for most of us but also an auction and by do if you want to see the output that's the report that we got from the research that's the one I will share with you so it's giving you both an executive summary and a ton of data points on how to use your tokens better yeah sorry last thing land graph land graph is basically a code if you
want to create a work rough with in land graph that's basically a code that's how it looks but another thing that is nice about land graph you can have a visualization built in done by land so even for those of you who are using code land graph has a built-in way to visualize the code that was created so even the top tier is not that scary in reality right a few more things before we maybe take a couple of questions the six habits that separate an effective agent orchestration or graph from a very expensive one some of those were mentioned there implicitly you want to match the another the model to the node part of the decision here is which model to use for which work or which type of work and of course you want to be cheap and fast for more mechanical steps and yes no verdicts and much more stronger models where judgment is needed each node ideally should get the relevant context only and what's being passed between different nodes is the contract so you need to be very careful and intentional about what passes between nodes maybe I draft an rubric
or finding in a format we you never need to pass the entire conversation that's not the right way to flow work for most cases the next thing I want you to spend wherever effication pays meaning that if you want to fan out to multiple agents there is a token cost to that if you need to summarize between nodes that can help if you can cap turns per node that can also help but be very deliberate about that because a beautiful graph like that can easily become a huge token consumer if you for example will use the built-in deep research by some of the tools those can easily take millions of tokens without bashing an eye so be careful about that I want you to verify early so add these nodes of verification and add clear boundaries whether it's because you're on a loop or just as part of the wall graph that you did because the especially the more complex your graph is the more compounding of mistakes become expensive and at least as of now humans are needed and for the most part superior than the beast so we'll have human at the right gate maybe it's at the end maybe it's
early on to approve the plan but humans should be intentionally brought into the graph one last thing that I want to cautious you again many people when they come to design a wall graph they think about exactly the way the work is being done by humans today and the way work is being done by human today is often highly bounded by human limitations attention span time bandwidths ability to be proficient in multiple things many of these limitations are not the relevant limitations for your agents and for your tools so I don't want you to just take the exact way that work is being offloaded between humans today and move it to a machine that's not the way to go that you need to understand either by testing or by understanding the limitations of the tools how to better configure the work given that in many cases these tools are much less prone to get confused or get tired than humans are and in many cases that requires a little bit of a radical thinking about
what's the bottom line what's the job to be done not what's the current processes of how humans do the work in order to design the best possible graphs out I'm sharing with that that we do later on but these are the concrete commands for the different tools that people are using so the concrete commands for loops and for using sub agents because things are moving so quickly always do a web search or consult with the tool attend as to what's the best way because sometimes they deprecate commands between you will wake one morning and the command is no longer there as well as different modalities for example the in the cloud code the sub task is only accessible in the CLI and not in the desktop so don't just assume that because there is a command that it's going to be operational in the self-specifically that you're using I think that if you're looking at loops versus graph then loops are more forgiving because the graph is a bit of a confession of how your work really flows who really owns what and work quality really gets decided and by the way this is why
you will probably be better and that than most engineers because you've spent your career learning how the specific work that you focus on and move through the organization and that's the knowledge that became the technical skill knowing how to inject your subject matter expertise into designing the these right the correct systems if I need to like summarize the hour on one slide and it's only relevant as of August 2026 because the thing will continue to grow probably by wintertime either with a new buzzword or with new skills as we get there but the best practitioners in knowledge work they've mastered all of these four they understand agents and they understand the underlying loop that is being implemented in the harness they know how to define concrete agent workflows they know how to configure loops to get the tools to unautonomously well and they know how to configure themes of agents that work effectively either as one task or in general to implement an entire organization that's the skill set for you to master as of August 2026 before I go to Q&A I will just do a quick plug
if you want to go much deeper depending on where you are we do have two training programs one is the executive ketchup for people who are a little bit let's call it behind all needs to make sure that they become best in class in AI usage and not just best effort and the executive agent leadership that's a much more advanced course for building themes of agents and configuring the strategy for your organization and so on and that way if do you want to add a few words I mean what more words are there I think the part of what makes this moment important is we've shifted from AI skills being useful new tools to actually being fundamental work primitive shifts and what I mean by that is that like when we started super intelligent a million years ago the very first iteration of the platform was like how to use mid-journey and how to prompt and and these things were valuable they were like nice skills to have they could get you leverage but the way that we did work hadn't
fundamentally changed yet we are now increasingly finding that big chunks of what we used to do instead our job is to now manage agents to do them and that is a transitional process it's not all at once it's not going to be all of the tasks that we do but we're all kind of involved now in the discovery to some extent of what it means to manage agents and so I think that that's the the lens through which I look at these things is we're all kind of like piece by piece giving ourselves an MBA in agent management and it's going to keep iterating and evolving but I think that a lot of these things the reasons that we cling on to loops and graph engineering some of these ideas is that they start to feel more like core primitives as opposed to just another fly-by-night skill or something like that so in the same way that you wouldn't expect yourself to know or to be perfect at advanced management techniques in a single session or experiment or a couple of days it's going to be the same with this it's just going to take hands-on work and experimentation and
there are no experts at this as I've said in the past there are just people who have done it more so even by virtue of being here I think you're probably ahead
More episodes
More from The AI Daily Brief: Artificial Intelligence News and Analysis

The Multiplayer AI Sprint: Build Your Team’s First Shared Agent
The AI Daily Brief: Artificial Intelligence News and Analysis

How to Build an AI-Native Company Today
The AI Daily Brief: Artificial Intelligence News and Analysis

How AI Changed This Summer
The AI Daily Brief: Artificial Intelligence News and Analysis

Why Fable 5.1 Is Worth the Upgrade
The AI Daily Brief: Artificial Intelligence News and Analysis