TL;DR
Why the discipline behind project controls may become essential to agentic enterprise work.
Enterprise AI is moving from answering questions to carrying out work. Microsoft’s introduction of a persistent agent that continues working when the user is absent is one recent sign of that change.
This creates a new management problem. If an AI agent keeps working across systems, decisions and handoffs, how does the organization know whether it is still advancing the intended outcome?
Logs can show what the agent did. Permissions can limit what it may access. Neither establishes the operational reference needed to judge progress, variance or acceptable change.
Project controls solved a version of this problem long ago through baselines. A baseline preserves the approved intent against which current performance can be understood. I believe the same principle will become essential for governing work performed by people and AI together.
Every consequential AI assignment should begin with an operational baseline: the outcome, owner, inputs, dependencies, constraints, approval points, evidence requirements and recovery path. Without that structure, agentic AI risks producing more activity than execution.

AI is beginning to work while we are away
For most of the generative AI era, the interaction has been familiar. A person asks a question. The model responds. The person decides what to do next.
That pattern is changing.
On 25 September, Microsoft introduced Autopilot as a persistent agent that can continue working when the user is absent. The announcement is part of a wider movement towards agents that monitor conditions, use tools, produce deliverables and advance work over time. Microsoft describes this as a new model of agentic work.
This is an important shift. It also exposes a control problem that becomes harder as AI becomes more capable.
A chatbot gives you an answer. A persistent agent creates a sequence of actions.
Once the work continues across hours, systems and organizational boundaries, the relevant questions change. What outcome is the agent pursuing? Which assumptions remain valid? What other work depends on its output? What evidence would justify completion? Which change requires a person to intervene? What happens when the operating conditions no longer match the original assignment?
An agent can execute every instruction correctly and still move the organization away from what it intended.
That is why I think enterprise AI needs to borrow an idea from project controls: the baseline.
What project controls understands about change
Anyone who has worked with Primavera P6 will understand why baselines matter.
A project schedule changes throughout execution. Activities move. Actual progress replaces forecasts. Constraints appear. Teams resequence work. An approved change may alter the plan itself.
Without a preserved reference, it becomes difficult to tell whether the project is advancing, drifting or merely being redrawn.
Oracle’s current P6 guidance describes the practical value clearly: comparing the current schedule with its baseline helps teams identify activities that start or finish later than planned and evaluate schedule and cost variance. The baseline provides the reference against which current conditions become meaningful.
A baseline does not freeze the project. It makes change governable.
That distinction matters far beyond scheduling. Every complex piece of work needs some durable expression of intent. People need to know what was approved, what has changed and whether the change is acceptable. When a new plan is adopted, the organization should preserve the history rather than quietly rewriting it.
AI agents need the same discipline.
An audit log is not an operational baseline
Many discussions about AI governance focus on access control, model risk, audit trails and human oversight. These are necessary. NIST’s AI Risk Management Framework, for example, calls for clearly defined human and AI responsibilities, documented oversight, deployment testing and continued monitoring of production behavior. Its framework connects governance with mapping, measurement and management throughout the AI lifecycle.
Yet a record of an agent’s actions cannot tell us whether those actions still serve the intended outcome.
Imagine an agent responsible for preparing a procurement package. Its activity log might show that it collected documents, compared vendor responses, created a recommendation and notified the commercial team. That is useful evidence.
The log does not necessarily reveal that the engineering specification changed midway through the work. It may not show that construction has resequenced the affected area, that the original decision deadline no longer applies or that a downstream activity now depends on a different technical assumption.
The agent may have completed its assigned steps. The work itself may no longer be valid.
An operational baseline would make that visible. It would preserve the accepted specification, required inputs, affected dependencies, accountable owner, decision deadline, approval route and evidence needed before the recommendation could advance.
The difference is simple.
A log tells us what happened.
A baseline helps us decide whether what happened still makes sense.

What an agent baseline should contain
An operational baseline does not need to become a vast specification that makes automation impossible. It needs enough structure to distinguish useful progress from uncontrolled motion.
For consequential work, I would expect it to define:
- the intended outcome and the accountable human owner;
- the exact activity or decision the agent is supporting;
- the inputs the agent may rely on and which systems remain authoritative;
- the assumptions and constraints that shape the assignment;
- the work, people and decisions that depend on the output;
- the actions the agent may take without further approval;
- the transitions that require human review;
- the evidence needed to accept the result;
- the conditions that should stop, reroute or invalidate the work;
- the recovery path if the agent fails or the operating context changes.
This structure creates a work contract around the agent.
It also allows organizations to use different models and specialist agents without making the model itself the operating system. The agent can change. The approved intent, dependencies, evidence and history remain part of the organization.
That may become one of the most important design choices in enterprise AI.

Human oversight should happen at transitions that matter
“Keep a human in the loop” sounds reassuring, but it is too vague to guide real operations.
Which human? Reviewing what? At which point? Against what evidence? With authority to approve which consequence?
If a person must inspect every minor action, the agent delivers little practical leverage. If human review appears only after the agent has made an external commitment, changed an authoritative record or triggered dependent work, the control arrives too late.
Project controls offers a better model. People do not approve every calculation in a schedule. They govern the baseline, review material variance, assess proposed changes and authorize consequential transitions.
Agentic work can follow the same pattern.
An agent may gather evidence, prepare an analysis or propose a change. A validation step checks whether the output meets the agreed contract. A named person reviews the consequential decision. An approval milestone then allows dependent work to proceed.
This gives human judgment a precise role. It also prevents an agent’s completion status from being mistaken for business approval.

Capital projects show why this matters
Capital projects make the problem unusually visible because the consequences of fragmented execution are expensive.
Recent Accenture research involving 1,050 leaders and frontline workers found a striking gap between leadership confidence and site experience. While 73 to 77 percent of senior leaders believed they positively influenced cost, planning and schedule performance, only 13 percent of workers said leadership decisions consistently shaped their daily work. Accenture argues that field signals are often delayed or filtered and that technology frequently reflects office priorities more closely than site reality. The report calls for an execution model that can turn site conditions into decisions quickly enough to change outcomes.
This is the current P6 discussion in practical terms.
P6 should remain authoritative for the controlled schedule. Teams still need a more detailed execution environment around that schedule: commitments, constraints, deliverables, evidence, decisions and handoffs. Changes in that environment should inform project controls through a governed feedback loop.
The same architecture applies when AI joins the team.
An agent might identify a missing deliverable, trace a constraint or prepare a status update. Its output should remain linked to the relevant work, source evidence and approval route. The scheduler retains authority over accepted schedule changes. The project manager retains responsibility for the operating decision. The agent improves the speed and quality of understanding.
This is a stronger model than allowing AI to rewrite the plan simply because it detected a variance.
The next operating advantage is controlled adaptation
The instinctive measure of AI progress is speed. How many tasks can an agent complete? How long can it work without supervision? How much human effort can it remove?
Those questions matter, but they do not measure whether the organization is becoming better at execution.
A more useful measure is controlled adaptation.
Can the organization detect when conditions have changed? Can it understand which work is affected? Can it revise the plan without losing the original reasoning? Can people distinguish an accepted change from an agent’s proposal? Does each completed assignment improve how the next one is structured and governed?
This is where baselines become more than a control mechanism. They create the foundation for learning.
When estimated effort, actual effort, accepted output, variance, rework and downstream effect remain connected, organizations can improve how they allocate work between people and agents. They can learn which agent performs well under which conditions. They can identify where human review adds value and where it merely adds delay. They can change the operating method while preserving the evidence behind that change.
The result is not unrestricted autonomy. It is increasing organizational capability.
Where Optimality fits
This is the direction we are pursuing at Optimality.
We see work as a connected execution network containing activities, people, agents, dependencies, decisions, deliverables, commitments and evidence. Existing enterprise systems retain authority for the records they govern. P6 can remain authoritative for the control schedule. Financial systems can retain posted actuals. Document systems can retain approved revisions.
Optimality provides the operating context around those records. It makes the assignment, execution, evidence, variance and human approval path visible as work moves between people and AI.
An activity can launch an AI Optimizer. Its output can return as connected content or a deliverable. A validation step can inspect the result. A person can review it. An approval milestone can govern what proceeds next.
That is what a baseline for agentic work begins to look like in practice.
The question leaders should ask now
Persistent agents will soon become ordinary parts of enterprise work. Some will monitor operations. Others will prepare decisions, coordinate information or carry tasks across several systems.
The leadership question is no longer limited to whether an agent can perform the work.
It is whether the organization has defined the reference against which that work can be trusted.
What outcome was approved? What assumptions support it? What changed? Which dependencies are affected? Who owns the decision? What evidence allows the work to advance? How can the organization recover when the agent or its context is wrong?
Project controls learned to ask these questions because complex work cannot be governed by activity alone.
The same lesson now applies to AI.
Every agent that carries consequential work should have a baseline. Without one, enterprises may gain enormous execution speed while losing the ability to tell whether they are still heading in the right direction.
Sources
Microsoft, “Introducing the new Copilot with Home, Code and Autopilot,” 25 September 2026
Oracle, “Comparing Current and Baseline Schedules,” Primavera P6 Professional User Guide Version 26
NIST, “AI Risk Management Framework Core”
Accenture, “The Capital Projects Execution Gap,” 5 August 2026





