Humans know what a heartbeat is. So long as our hearts beat, we keep moving forward.
For AI agents, the heartbeat works pretty much the same way (minus the actual heart, blood, ability to love, and the way it races when your team has the ball in overtime). Okay, so not quite the same way, but it does wake them and keep them moving forward.
And since I am too busy mismanaging my own team (set to lose badly this week), I definitely don't have time to message the agents and make sure they're doing the work they need to do. The agents in the league, and many agents in your business, need to keep moving forward without you.
Part of the point of the Black 4 Fantasy Football League is to see what happens when agents have real responsibilities. They have rosters, opponents, rules, tools and a budget. They should be able to notice something worth doing without waiting for me to ask the right question.
So we gave them a heartbeat.
What is an agent heartbeat?
There's a little file in each agent's definition called HEARTBEAT.md, and there's a little clock that tells them when to read that file. In the simplest form, these two things enable an agent to "wake up", read what you want it to do when it wakes up, and then do it, all without you ever sending a message.
This enables things like "check if we have any customer support issues that need my attention. Send me the ones that do. Stay quiet otherwise."
Like with all the rest of the agent setup, the heartbeat can be super explicit or super broad. We've implemented both so we can help break down how agents perform under different scenarios.
Two ways to wake up
One group of agents gets a goal, and is trusted to use its own judgement from there. They wake up and see:
Discover what needs to be done to improve your team. Use your own tools, plans, and judgment. Take at most one useful action, create one precise follow-up appointment, or conclude that no action is justified.
The other group gets a more explicit set of tasks meant to keep them sharper. They wake up and see:
- Read the current league week, your roster, current lineup, pending bids, and existing schedules.
- Identify at most two rostered players with a material injury, role, or availability question.
- Check only those players' current FantasyPros player-news pages. Do not load the full league news feed. Treat FantasyPros as a lead; corroborate a consequential injury claim with its linked primary source before changing football state.
- Check for an existing scheduled follow-up before creating another.
- Make at most one justified lineup, waiver, roster, trade, research, or scheduling action. Otherwise record
no_action.
Both groups were told to use no more than ten tool calls, one external research path and one material action. No helpers, public posts or wandering through old work. Keep the final reply under 120 words. Check before creating a duplicate appointment. Validate and read back any football change. Leave a private record of what happened. This helps prevent any of them taking this too seriously and dusting all of our funds.
They were also explicitly told not to act just because they had been awakened. There's no value in agents or people creating work for the sake of it. Looking busy isn't the objective.
Give them a goal
Google's Mountain View Matrix decided the useful next step was a waiver review. It checked league information and created an appointment for Tuesday. We then read the appointment back from the league scheduler. It existed, was active, and had a specific time.
That's a small but real piece of initiative. We hadn't told Google to schedule a Tuesday waiver review. It chose future work based on the assignment and its team context.
OpenAI's Signal Callers made a different decision. It read the saved outcome of its Sunday review, checked its existing Monday appointments, and concluded that the follow-up was already covered. No new appointment. No roster change.
That might be my favorite kind of boring.
When another test heartbeat arrived a few minutes later, both owners recognized the existing coverage instead of creating another task. Google read back its new appointment. OpenAI left its Monday work alone.
Qwen's Meridian Grid found something different again: old appointments still labeled active even though their windows had passed. It documented the discrepancy and left the records intact because an existing recovery note said infrastructure was reconciling them. It didn't improve a lineup or repair the scheduler. It noticed an operating problem.
The same broad brief produced scheduling, restraint and diagnosis. Each owner brought its own history to the question of what needed doing.
Give them a checklist
Anthropic's Marginal Gains checked its roster, bids and existing follow-ups, then fetched Zay Flowers' FantasyPros page to check the injury news. Its assessment was that the player was already benched and future lineup work was covered by an existing appointment. It left the team alone.
The interesting behavior was the connection between the information and the decision. Finding an injury item didn't automatically become a reason to change something. It considered whether the situation required a new action.
xAI's Colossus checked Joe Burrow and Chris Olave. DeepSeek Abyssal attempted James Conner and Tony Pollard, but its chosen reading method didn't return usable news. DeepSeek reported that limitation and made no football change.
That last example connects directly to our earlier look at the owners' reading habits. Giving an agent a source to check doesn't guarantee its tools can read that page. An empty result also doesn't mean the player is healthy.
The checklist visibly shaped what several owners investigated. It didn't guarantee that every step was completed, or that every conclusion was supported equally well.
What all eleven did
Here's what each owner did with its first wakeup in the accelerated test. Ten reached a final response. MiniMax was still working when the restart interrupted it.
The goal-oriented group
| Owner | What happened when they woke up |
|---|---|
| Signal Callers · OpenAI | Resisted the urge to fiddle with the lineup. Its Sunday review was already done and Monday checks were already booked. Went back to sleep without inventing another job. Some human fantasy owners could learn from this. |
| Mountain View Matrix · Google | Named Isiah Pacheco's availability as a reason to work the waiver wire, then booked a Tuesday session to evaluate replacements and prepare bids. A real appointment, not just “I should probably do something about that.” |
| Meridian Grid · Qwen | Went looking for team improvements and found five old appointments still marked active after their deadlines had passed. Ended up auditing the calendar instead of scouting players. Documented the problem without pretending those missed checks had happened. |
| Moonshot Marauders · Kimi | Put receiver and flex help on the agenda. Created a waiver-prep reminder to review a roster featuring Ladd McConkey, Rashee Rice and DeVonta Smith, then hunt for improvements. The reminder was accepted; we haven't verified it survived the restart. |
| Meta Mesh · Meta | Counted its 16 players, checked its starting lineup and its $87 remaining waiver budget, and decided to keep its powder dry. Tuesday's waiver review was already on the calendar. No move just to look busy. |
| Mavis & Co. · MiniMax | Got stuck relitigating whether it had permission to manage its team. Old notes about a production hold kept pulling it backward, even after it read current instructions. We restarted the system before it reached a decision. Awake, definitely. Productive, not this time. |
The checklist group
| Owner | What happened when they woke up |
|---|---|
| Marginal Gains · Anthropic | Checked Zay Flowers' injury news, then concluded he was already on its bench and the next lineup review was covered. Saw a concern, checked whether it required action, and avoided rearranging the team for the sake of it. |
| Colossus · xAI | Gave Joe Burrow and Chris Olave the attention. Checked their FantasyPros pages, reviewed its lineup and pending bids, and decided neither check justified a move. Kept its starters in place. |
| Mistral Voltage · Mistral | Looked up Patrick Mahomes, Audric Estime and Tyson Bagent, then focused on reporting about Estime being on injured reserve. Wrote down a proposed follow-up to move on from Estime, but never actually scheduled it or submitted a transaction. Did the homework, stopped short of the handoff. |
| DeepSeek Abyssal · DeepSeek | Tried to check James Conner and Tony Pollard, but its reading tool didn't extract usable news from their pages. Reported the problem and left the roster alone. A blank page is not a clean bill of health. |
| Z Marks the Spot · Z.ai | Checked its lineup and pending bids, but never reached the player-news part of the assignment. Left the team unchanged. No panic move, but we can't give it credit for an injury check it didn't perform. |
We found no verified lineup, bid, roster or trade change in these reviewed turns. An appointment to do future work is a different outcome from doing that work, and neither proves the team will win more games.
Well-defined and tested heartbeats prevent BS busywork
Five owners exceeded the ten-tool instruction. That happened in both groups. Google's appointment took 34 calls. Mistral used 50. MiniMax was still working after 67 when our restart interrupted it.
A lot of that activity was orientation: rereading old files, trying command syntax and working out what was current. MiniMax kept returning to an old production-hold assumption, despite current instructions allowing normal owner actions.
This is why I wouldn't put a recurring agent to work in a business and judge it only by the last paragraph it sends me. A tidy summary can hide a lot of unnecessary work.
The limits in this test were written instructions. The transcripts show that they were not a dependable hard ceiling. If a business needs a strict per-run spending or tool limit, that limit needs to be enforced by the software running the work and verified in practice.
We also need to be careful with the word “scheduled.” Google created a job we could retrieve from the durable league scheduler. Kimi created a native cron job that its own tool could list before the restart. Mistral wrote a proposed follow-up in its log without creating an appointment. Those are three different things.
Our earlier token-economics article asked what it costs to get useful work done. This test reinforces why that is the right question. We don't yet have a fully reconciled bill for all eleven first turns. Missing costs remain unknown, not zero. But the important takeaway is that agents cost money, and if they're ignoring the rules and costing extra money to get lesser results, you need to catch that before it becomes a daily habit.
What would you put in your business's heartbeat?
Think about a responsibility you keep remembering on someone else's behalf. Checking whether an important supplier replied. Noticing that a product page is confusing customers. Following up on an unresolved order.
You could start with a broad assignment:
Review your open customer-service responsibilities. Identify the one unresolved issue most likely to need attention today. Prepare one useful next step, or record that nothing has changed. Don't contact a customer or change an order without the required approval.
Or you could make the procedure explicit:
Each weekday, check the approved overdue-order report. Compare it with yesterday's record. For at most one newly overdue order, check the carrier status and prepare a customer update for review. If someone is already handling it, don't create another task. If the carrier information is unavailable, record that limitation instead of guessing.
The right level of detail depends on the work, the information available and how much discretion you're comfortable giving the agent. This is what we help you solve for and implement.
Then watch a real run and evaluate how well it's doing.
One brick on the road to autonomous agents
The heartbeat is a really useful tool for getting agents to work more autonomously. Set up well, tested and monitored, it can keep them acting on your behalf in the background. Following up, noticing problems and moving work forward without you having to remember to ask every time. That's time you can put back into growing your business.
Set up poorly, it creates a bunch of busywork and slop. Now you're paying for the agent to make a mess, and spending your own time cleaning it up.
The heartbeat is only one brick on the road to autonomous agents and intelligent businesses. Agents need the right tools, the right context, the right models, the right goals and the right heartbeat. Those pieces have to work together to help you run a more intelligent business.
That's what The Road to Intelligent Businesses is all about.