
Model-Agnostic AI for Government Sales
Why single-model AI and shared public data will raise costs and erase differentiation for businesses selling to government.
A recent preregistered AI study offers an important warning for business leaders. Participants often preferred an AI system that validated their views, not because its advice was better, but because it made them feel understood. The enterprise lesson is not about personal relationships. It is about measurement. Users can prefer an AI experience that feels productive without receiving a better result. In business software, more conversations, longer responses, repeated revisions, and higher agent activity can create the appearance of adoption while consuming more tokens and producing no corresponding improvement in decision quality.
Businesses selling to government are particularly exposed to this problem. Many are adopting the same basic AI strategy: select a leading model, connect it to public procurement data, add a conversational interface, and deploy agents around existing sales processes.
This approach can improve baseline productivity. It is unlikely to create a durable advantage.
When competitors use the same model families and the same public data, their outputs will increasingly converge. When every task is forced through one general-purpose model, operating costs become less predictable. When the model also controls the workflow, the company becomes dependent on a vendor whose pricing, behavior, availability, and relative performance can change faster than a government procurement cycle.
So the real durable advantage will come from the operating layer that determines how models, data, software, and people work together to produce a trusted outcome.
Cheaper tokens do not guarantee cheaper AI
The economics of AI are commonly evaluated through the price of a model or the cost per million tokens. Both are incomplete measures.
Stanford’s 2025 AI Index found that the inference cost of a system performing at roughly the level of GPT-3.5 fell more than 280-fold between November 2022 and October 2024. It also found that open-weight models narrowed the performance gap with closed models from 8 percent to 1.7 percent on some benchmarks in a single year. The cost of accessing intelligence is falling, and the number of viable model choices is growing.
But the price per token is only one side of the equation. The other is how many tokens a system consumes before it produces a usable result.
A 2026 preprint examining agentic software-development workloads found that agent tasks consumed 1,000 times more tokens than code-reasoning and code-chat workloads. Runs on the same task varied by as much as 30 times, and higher token usage did not consistently produce higher accuracy. The models also struggled to predict their own consumption and systematically underestimated their eventual costs. These findings come from coding workloads, so the exact numbers should not be generalized to every enterprise process. The broader economic pattern is still important: agent costs are heavily influenced by context, architecture, retries, and stopping conditions, not merely the model’s published price.
Tool use introduces another layer of cost. Research published in the Findings of ACL 2026 evaluated Model Context Protocol configurations across 20 servers and 169 tools. In systems using customized clients, 56 to 72 percent of tokens were spent on planning and injecting tool schemas into the model’s context. Actual tool execution accounted for a negligible share of total cost. In other words, much of the expense came from telling the model what it could do, not from doing the work.
Consider what this means for a government opportunity or proposal workflow. A poorly designed agent might repeatedly load an entire solicitation, its addenda, the seller’s product documentation, prior proposals, customer references, CRM history, and dozens of tool definitions. It might then use the same frontier model to extract deadlines, classify requirements, compare references, draft responses, review its own draft, and revise it several times.
Every stage can replay much of the same context. Every agent may create another planning cycle. Every retry can trigger another full-model call. The system may appear active and intelligent while the underlying cost per completed proposal continues to rise.
For a company selling AI under a fixed subscription, uncontrolled agent activity becomes a direct gross-margin risk. Increased usage may look like successful adoption while quietly increasing the cost of goods sold.
The relevant business metric is not cost per token. It is cost per trusted outcome: the cost to accurately qualify an opportunity, produce an accepted compliance matrix, generate a draft that survives human review, or complete an auditable deliverable.
A single-model strategy creates concentration risk
The model market is moving too quickly for any business to assume that today’s leader will remain the best choice for every task.
Stanford’s 2026 AI Index reports that U.S. and Chinese models have traded the performance lead multiple times since early 2025. It also describes a “jagged frontier” in which models can solve competition-level mathematics while failing much simpler perceptual tasks. Agents have improved substantially on computer-use benchmarks, but still fail approximately one in three attempts on structured tasks.
A model that is strongest at complex synthesis may not be the most accurate or economical option for entity matching. The best model for extracting requirements may not be the best one for evaluating product fit. A model that performs well on proposal writing may perform poorly when it must follow a rigid approval policy or update several business systems in sequence.
A single-model architecture does not eliminate this complexity. It concentrates it inside one external dependency.
That creates several forms of exposure:
The model provider may change pricing, usage limits, or commercial terms.
A new model version may behave differently from the version previously tested.
Safety policies or system behavior may change without corresponding changes to the customer’s workflow.
An outage or capacity constraint may stop every AI-supported process.
A better model may emerge, but switching could require rebuilding prompts, tools, evaluations, and integrations.
The risk is amplified in government markets because procurement and implementation can outlast several generations of model development. In April 2026, the Government Accountability Office reported that some agency officials described federal acquisition processes that can take up to two years. GAO also warned that rapid technology development can render infrastructure and business cases obsolete, and noted that vendor-driven AI purchases may create unanticipated cost growth and declining model reliability. One GSA official compared the current market to the early search-engine era, when buyers could easily have committed to a provider that was later displaced.
Government guidance is already moving toward portability. Current federal AI acquisition policy emphasizes competition, performance-based purchasing, clear requirements, and avoidance of vendor lock-in. OMB guidance summarized by GAO recommends that agencies seek data and model portability, clear licensing terms, pricing transparency, ongoing evaluation, and contract provisions addressing vendor lock-in.
Businesses selling AI to government should expect their architectures to face the same questions:
Can the model be replaced without losing the workflow? Can the customer’s data and operating knowledge be moved? Can the system compare model performance over time? Can costs be explained at the task and outcome level? Can the vendor prove what information produced an answer?
A product hardwired to a single model will have increasingly weak answers.
Shared public data produces parity, not differentiation
The second strategic mistake is treating access to public data as a durable competitive advantage.
At the federal level, anyone can search SAM.gov contract opportunities without an account. Opportunity data can also be downloaded or accessed through a public API. The USAspending API provides public access to comprehensive federal spending and award data, including recipient, agency, and geographic information.
These are valuable sources. They are not rare sources.
The same applies whenever competitors monitor the same public budgets, meeting agendas, legislative records, procurement forecasts, solicitations, award notices, contract records, and agency websites. AI makes these sources easier to collect, normalize, search, summarize, and present. As those capabilities become broadly available, data aggregation becomes table stakes.
Classic resource-based strategy holds that sustained advantage comes from resources that are valuable, rare, difficult to imitate, and difficult to substitute. Public data may be valuable, but by definition it is not rare or difficult for competitors to acquire.
This does not make public data unimportant. It changes where value is created.
A meeting transcript is not automatically intent. A budget line is not automatically a qualified opportunity. A new executive hire is not automatically a buying signal. An RFP is not automatically evidence that a company is well positioned to win.
The value comes from interpreting a public event against private operating context:
Does the requirement match something the company can actually deliver?
Does the seller have relevant customers who will serve as references?
Is there credible evidence that the product can solve the stated problem?
Does the company have the contract vehicle, partner relationship, geographic coverage, or implementation capacity required?
Has the account considered the seller before?
Is the opportunity early enough to influence, or have the requirements already been shaped?
What did the company learn from similar pursuits, losses, implementations, and renewals?
Two companies can receive the same public signal and reach very different conclusions because their products, relationships, evidence, credibility, capacity, and risk tolerance are different.
That proprietary context is where differentiation begins.
In government sales, the public opportunity is often not the beginning
Government sales intelligence tends to overvalue the visible procurement event.
At the federal level, agencies are required to conduct market research before developing new requirements documents and before soliciting offers above applicable thresholds. Current FAR guidance also encourages responsible and constructive exchanges with industry during market research.
A published solicitation can therefore represent a mature procurement event, not the beginning of the buying process. By the time an RFP appears, the agency may already understand the available solutions, have defined its requirements, selected a procurement strategy, established evaluation criteria, and developed expectations based on earlier market engagement.
This is why giving every seller faster access to the same RFP does not necessarily change competitive outcomes. It may simply help more companies respond to an opportunity after their probability of winning has already diverged.
The higher-value questions are upstream:
Which accounts are entering a budgeting or market-research process? Where does the seller have the credibility to engage early? Which references will matter to the buyer? What evidence will reduce the perceived risk of selecting the company? What action must happen before the requirements become fixed?
Public data can help answer parts of these questions. It cannot answer them without company-specific context.
Public data is raw material. The operating layer is the moat.
Orchestration can matter more than model size
The alternative to a single-model strategy is not simply adding more models. An unmanaged collection of models can increase complexity and cost.
The answer is structured orchestration.
RouteLLM, presented at ICLR 2025, demonstrated that a routing system could dynamically select between stronger and weaker language models. On public benchmarks, the approach reduced costs by more than two times without sacrificing response quality, and the routing logic continued to perform when the underlying model pair changed.
A separate 2026 benchmark examined AI agents following business procedures. Researchers represented the required workflow as a state machine, exposed only the tools needed at each stage, and replaced old instructions as the process advanced. Under this architecture, a smaller model achieved a higher average business-adherence score than a larger model using a static prompt, 0.649 compared with 0.564. The structured workflow contributed more to the outcome than simply selecting the larger model.
The lesson is significant:
The strongest model used poorly can lose to a smaller model embedded in a better workflow.
For a government revenue process, the operating layer should determine how each part of the work is performed. Deterministic software can handle date calculations, required-field checks, deduplication, formatting, and rule enforcement. Smaller or open-weight models can perform bounded, high-volume extraction and classification tasks. Frontier closed models can be reserved for ambiguous reasoning, evidence comparison, and high-value synthesis. Separate validation steps can check citations, completeness, policy adherence, and numerical consistency. Human approvals can govern consequential decisions such as qualification, pricing, commitments, and final submission.
Open-weight models are not automatically superior, and closed models are not automatically too expensive. The correct choice depends on the task, data sensitivity, required accuracy, latency, deployment constraints, and economic value of the outcome.
The goal is not to pretend that models are interchangeable. They are not.
The goal is to make the business independent of any one model while remaining highly selective about which model performs each task.
What a durable AI operating layer must own
A durable operating layer should keep the company’s business logic outside the model. It should own five critical controls.
1. Work decomposition and state
The layer should divide a complex process into bounded tasks, maintain the current workflow state, and define what must happen next. The model should not be responsible for remembering the entire process or deciding indefinitely whether the task is complete.
2. Context and evidence
The layer should retrieve only the information needed for the current task. It should manage permissions, provenance, document versions, customer references, product evidence, and account history without repeatedly loading the entire company knowledge base.
3. Model routing and economic controls
Each task should be matched to the appropriate model, software function, or human. The layer should enforce token budgets, retry limits, latency thresholds, fallback routes, and stopping conditions. A failed extraction step should not require rerunning an entire proposal workflow.
4. Evaluation and governance
Outputs should be tested against explicit requirements. The layer should preserve citations, confidence indicators, approval records, tool activity, model versions, and the reasoning evidence necessary for audit and review.
5. Outcome feedback
The system should learn from what happened after the output was created. Was the opportunity accepted or disqualified? Did the proposal pass review? Was the requirement interpretation correct? Did the company win? Did the customer become a credible reference? That feedback should improve routing, qualification criteria, evidence selection, and workflow performance.
These controls create more than model portability. They create institutional memory.
The architecture Civio is building
At Civio, we are building an operating layer for businesses selling to government based on this premise, and we will soon be introducing this same approach directly to state and local governments.

The system does not ask one model to do everything. It breaks revenue work into governed tasks across account intelligence, qualification, pursuit, proposals, delivery, and growth. Public signals are evaluated against the company’s private operating context, including its products, services, customer evidence, references, historical opportunities, delivery capabilities, relationships, and internal decision criteria.
Each task can then be directed to the most appropriate interaction: deterministic software, a bounded open-weight model, a frontier closed model, an independent validation process, or a human approval.
The workflow state, business rules, source material, permissions, approval history, and outcome data remain in the operating layer. The underlying model can change without requiring the company to rebuild how it qualifies opportunities, develops proposals, governs commitments, or learns from results.
This architecture also changes how AI performance should be measured. Instead of maximizing messages, prompts, agent runs, or total usage, the system can optimize for cost per opportunity evaluated, cost per accepted draft, percentage of requirements supported by evidence, human review time, retry frequency, and the reliability of the completed outcome.
Companies that rely on the same public data and the same single model may achieve temporary feature parity. They will not create much competitive separation, and they may discover that growing AI usage produces growing costs without improving win rates or customer outcomes.
The durable asset is the company-specific context encoded in governed workflows, strengthened by evidence, and improved through real operating results.
The model will keep changing. The operating layer that knows how the company finds, qualifies, wins, and delivers government work should not.
Sources
arXiv, “Social impacts of sycophantic AI.” Preregistered research examining user preference for validating AI systems and the distinction between perceived support and objective usefulness.
https://arxiv.org/abs/2605.07912Stanford Institute for Human-Centered AI, AI Index Report 2025. Data on declining inference costs, model performance, and narrowing differences between open-weight and closed models.
https://hai.stanford.edu/ai-index/2025-ai-index-reportarXiv, research on token consumption in agentic software-development workloads. Analysis of token variability, agent consumption, cost predictability, and the relationship between token use and accuracy.
https://arxiv.org/abs/2604.22750ACL Anthology, Findings of ACL 2026. Research evaluating Model Context Protocol implementations, tool schemas, planning overhead, and token consumption.
https://aclanthology.org/2026.findings-acl.1967/Stanford Institute for Human-Centered AI, AI Index Report 2026. Analysis of changes in model leadership, uneven model capabilities, and progress in agent performance.
https://hai.stanford.edu/ai-index/2026-ai-index-reportU.S. Government Accountability Office, Artificial Intelligence Acquisition. Analysis of government AI procurement challenges, acquisition timelines, vendor dependence, technology obsolescence, and cost risk.
https://files.gao.gov/reports/GAO-26-107859/index.htmlThe White House, Federal AI Use and Procurement Guidance. Federal policy emphasizing competition, performance, portability, pricing transparency, and avoidance of vendor lock-in.
https://www.whitehouse.gov/fact-sheets/2025/04/fact-sheet-eliminating-barriers-for-federal-artificial-intelligence-use-and-procurement/SAM.gov and USAspending.gov. Public federal procurement, contract opportunity, award, and spending datasets demonstrating the broad availability of government-market information.
https://sam.gov/opportunities
https://www.usaspending.gov/Jay Barney, “Firm Resources and Sustained Competitive Advantage,” Journal of Management. Foundational resource-based strategy research describing the characteristics required for resources to create sustained competitive advantage.
https://journals.sagepub.com/doi/10.1177/014920639101700108Acquisition.gov, Federal Acquisition Regulation Part 10. Federal requirements and guidance concerning market research and exchanges with industry before solicitation.
https://www.acquisition.gov/far-overhaul/far-part-deviation-guide/far-overhaul-part-10RouteLLM, ICLR 2025. Research demonstrating dynamic routing between language models to reduce costs while maintaining response quality.
https://proceedings.iclr.cc/paper_files/paper/2025/hash/5503a7c69d48a2f86fc00b3dc09de686-Abstract-Conference.htmlarXiv, research on AI agents following structured business procedures. Evaluation showing how workflow state, constrained tool availability, and structured orchestration can improve agent adherence and allow smaller models to outperform larger models using less structured approaches.
https://arxiv.org/html/2601.00596
About the Author
James Ha is the CEO and Co-Founder of Civio, an AI-native operating layer for businesses selling to government. He has spent more than 25 years building, scaling, and leading technology companies, including nearly two decades in businesses serving the public sector. His experience spans startups, high-growth software companies, acquisitions, and successful exits. James also advises founders, investors, and private equity firms on growth, strategy, and the changing role of AI in business. His work focuses on how technology can simplify the way businesses and governments work together.




