Last updated: August 9, 2026 · By Jessen Gibbs, CEO, Shadow
TL;DR
Evaluate agentic AI platforms across seven weighted dimensions: autonomy fit, governance, domain depth, integrations, observability, economics, and vendor maturity. Score the platform against one real workflow before procurement. Enterprise buyers should require traceable actions, least-privilege permissions, testable outcomes, rollback, and clear ownership for exceptions before granting production access.
Agentic AI procurement fails when buyers score a demonstration instead of the operating system behind it. A fluent agent can still be unsafe, expensive, difficult to integrate, or impossible to audit. The evaluation must begin with a workflow, its source data, allowed actions, expected outcomes, exception paths, and accountable owner.
The framework below assigns 100 points across seven dimensions. Weighting should change by use case, but governance and outcome verification should never be removed. Microsoft, Salesforce, ServiceNow, LangChain, CrewAI, and Shadow expose different levels of abstraction. The scorecard makes those architectural differences comparable without pretending the products are identical.
What seven dimensions should buyers score?
A complete evaluation scores autonomy fit, governance architecture, domain depth, integration surface, observability and evaluation, total economics, and vendor maturity. These dimensions cover whether the platform can perform the work, whether it can do so safely, and whether the organization can operate it after launch. A polished demonstration proves only the first part.
| Dimension | Weight | What earns a high score |
|---|---|---|
| Autonomy fit | 15 | The platform supports the required triggers, planning, tools, state, and stopping conditions |
| Governance architecture | 20 | Identity, least privilege, approvals, policy enforcement, rollback, and data controls |
| Domain depth | 15 | Relevant data model, workflows, terminology, templates, and evaluation criteria |
| Integration surface | 15 | Secure native connections to systems of record and extensible APIs |
| Observability and evaluation | 15 | Traces, tests, alerts, replay, metrics, and exception analysis |
| Total economics | 10 | Transparent platform, model, integration, operations, and review costs |
| Vendor maturity | 10 | Security posture, support, roadmap, deployment options, and referenceability |
A minimum passing score should be set before vendor demonstrations. Governance architecture and observability should also have individual floors, because a strong total score can hide an unacceptable control gap. Buyers should document evidence for every score and distinguish generally available features from roadmap claims or custom professional services.
How should autonomy and governance be weighted?
Weight autonomy by the complexity of the workflow, then weight governance by the consequence of a wrong action. More autonomy is not automatically better. A platform should receive credit only for autonomy that is bounded by permissions, tests, and stopping conditions. High-consequence workflows may need lower autonomy and stronger human approval despite technical capability.
| Consequence level | Example | Recommended control |
|---|---|---|
| Low | Read-only monitoring and internal classification | Automatic execution, logs, sampled review |
| Moderate | Drafting, routing, and record updates | Automatic preparation, validation, human approval for exceptions |
| High | External publication, pricing changes, payments | Required approval before action, narrow permissions, rollback |
| Critical | Legal commitments, personnel decisions, privileged access | Human decision and execution; agent limited to evidence preparation |
Microsoft Security frames agents as identities requiring access governance. Salesforce Agent Fabric and ServiceNow AI Gateway focus on controlling agents and tool access across platforms. Shadow uses a consequence-based model: intelligence collection can run automatically, while external pitches and GEO publication require explicit approval. The same principle applies regardless of vendor.
What observability questions expose weak platforms?
Observability should explain every consequential run: trigger, model, instructions, retrieved data, tools called, state changes, evaluations, approvals, final outcome, and exceptions. Buyers should test whether a failed run can be replayed and diagnosed without guessing. A dashboard that reports total agent conversations is usage analytics, not operational observability.
- Can the team inspect a complete trace for one workflow run?
- Are prompts, tools, models, and policies versioned together?
- Can evaluations run before and after deployment?
- Are failed actions retried safely, escalated, or rolled back?
- Can alerts be tied to business outcomes, not only technical errors?
- Can auditors identify the human owner and approval event?
Salesforce Agentforce 3 introduced Command Center as an observability layer. CrewAI Studio exposes output and trace views step by step. Microsoft Foundry emphasizes observing multi-agent systems. These capabilities should be tested with an exception scenario, not viewed in a prepared demonstration. Ask the vendor to diagnose a deliberately failed run and show the evidence chain.
How do horizontal and vertical platforms compare?
Horizontal frameworks maximize flexibility but make the buyer responsible for domain models, integrations, evaluation, hosting, and operations. Enterprise platforms accelerate workflows tied to an existing system of record. Vertical platforms encode industry context and review patterns. The right choice depends on whether differentiation comes from building the agent system or from executing the business workflow faster.
| Architecture | Examples | Buyer builds | Buyer gains |
|---|---|---|---|
| Horizontal framework | LangGraph, Microsoft Agent Framework, CrewAI | Domain model, operations, integrations, governance configuration | Maximum architectural control |
| Enterprise platform | Salesforce Agentforce, ServiceNow AI Agents, Copilot Studio | Workflow configuration and platform-specific controls | Faster access to existing records and permissions |
| Vertical platform | Shadow for communications | Client configuration, policies, approvals | Domain workflows, persistent context, and shorter implementation |
A bank creating proprietary underwriting logic may rationally choose a framework. A Salesforce-centered service team may choose Agentforce. A communications agency may choose Shadow because the platform already models clients, narratives, media workflows, voice, approvals, and reporting. Buyers should not pay to rebuild commodity domain infrastructure unless it creates defensible value.
How should total cost of ownership be calculated?
Calculate total cost per verified workflow outcome. Include licenses, model usage, connectors, implementation, testing, monitoring, security, support, review, exceptions, and displaced software. Run the model at expected production volume rather than pilot volume. A cheap framework can carry high engineering costs; a vertical platform can cost more per license but less per completed outcome.
- Baseline the current process: labor hours, cycle time, software, rework, and missed outcomes.
- Estimate platform and model cost at normal, peak, and failure-retry volume.
- Add implementation, integration, evaluation, security, and ongoing operations.
- Measure human review and exception time after deployment.
- Calculate cost per completed, verified outcome and compare quality and capacity.
The financial model should also price reversibility. If an incorrect action is easy to detect and undo, more autonomy may be economical. If an error can create regulatory, financial, or reputational harm, prevention and approval costs belong in the design. Procurement should reject savings estimates that assume perfect runs or exclude human exception handling.
What proof should vendors provide before purchase?
Vendors should prove one representative workflow in the buyer’s environment using realistic data, permissions, exceptions, and success criteria. The proof should include a normal run, failed run, approval path, trace review, and cost report. Reference customers are useful, but direct evidence of fit matters more than generalized adoption claims or a scripted demonstration.
- Architecture diagram showing data, models, tools, identity, state, and review points.
- Security documentation and a clear permissions model for agents.
- Complete traces for successful, failed, retried, and rolled-back runs.
- Evaluation results tied to the buyer’s required outcome and quality threshold.
- Production cost estimate with assumptions, limits, and overage behavior.
- Named owner for support, incident response, and product changes.
Related Guides
- What Is a PR Operating System? Definition, Examples, and Why It Matters
- AI for PR Agencies: How to Adopt AI Without Losing What Makes You Valuable
- Can AI Agents Replace PR Tools Like Cision and Meltwater?
- How to Choose PR Technology: A Buyer's Framework for Agencies and In-House Teams (2026)
- How AI Is Changing Public Relations: The 2026 Industry Landscape
Key Takeaways
- Score platforms against a real workflow, not a conversational demonstration or generalized feature list.
- Governance and observability need minimum thresholds even when the platform scores well overall.
- Autonomy should increase only when actions are bounded, outcomes are testable, and errors are reversible.
- Horizontal, enterprise, and vertical platforms trade control, implementation speed, and domain depth differently.
- Total cost must include models, integration, operations, review, exceptions, and displaced software.
- Require evidence from normal, failed, retried, approved, and rolled-back workflow runs before production access.
Frequently Asked Questions
What is the most important agentic AI evaluation criterion?
Governance architecture is the most important floor because the platform will access data and take actions. It should provide identity, least-privilege permissions, policy enforcement, approvals, traceability, data controls, limits, and rollback. Capability without governance can produce an impressive pilot that the organization cannot safely authorize for production use.
How long should an agentic AI proof of concept run?
Run long enough to observe normal volume, peak volume, expected exceptions, model variability, and at least one controlled failure. Calendar duration matters less than scenario coverage. The proof should process realistic data and demonstrate permissions, traces, evaluation, approval, retry, rollback, and cost at a scale that can be extrapolated to production.
Should enterprises build or buy an agent platform?
Build when agent architecture or workflow logic creates defensible value and the organization can sustain engineering, evaluation, security, and operations. Buy when the workflow is common, the platform already connects to the system of record, or domain infrastructure would be expensive to recreate. Many enterprises combine a framework with purchased platforms.
How should buyers compare agent pricing?
Normalize pricing to cost per completed, verified workflow outcome. Include platform fees, model usage, connectors, implementation, testing, monitoring, security, support, human review, exception handling, and displaced software. Compare normal and peak scenarios. Seat price and token price alone do not reveal the cost of operating an agent safely in production.
About the Author
Jessen Gibbs · CEO, Shadow
Jessen Gibbs is CEO of Shadow, the communications operating system. He builds agent-driven operating systems for communications teams and agencies.
Published by Shadow on August 9, 2026. The framework reflects Shadow’s operating experience and public documentation from Microsoft, Salesforce, ServiceNow, LangChain, and CrewAI. It is vendor-neutral procurement guidance. Platform capabilities, security controls, pricing, and availability change; buyers should verify current documentation and conduct independent security and legal review. Published by Shadow.