Sapien’s Trevor Koverko

Nearly four years after ChatGPT brought artificial intelligence into the mainstream, the industry is confronting a new constraint.

For years, progress was measured largely through scale: more parameters, more computing power, and more training data. That formula produced major gains. But as businesses move from experimenting with chatbots to deploying AI agents inside real workflows, model capability is no longer the only issue. Companies also need to know whether the data, evaluations, and outputs behind those systems can be trusted.

Trevor Koverko, co-founder and chief strategy officer of Sapien, believes verification will become a core part of the AI stack.

"Scale got AI to this point, but scale alone will not get it through the next phase," Koverko said. "When AI starts making decisions and taking actions, companies need more than a capable model. They need evidence that the data and judgment behind the decision are reliable."

Koverko's role at Sapien is focused on strategy, partnerships, market positioning, and enterprise adoption. He is helping turn the company's Proof of Quality protocol into a product that AI developers, data providers, and enterprises can use without replacing their existing systems.

That work follows a pattern from his earlier companies. Koverko founded Polymath, which helped create infrastructure for regulated digital securities, and was involved in building Polymesh, a blockchain designed specifically for regulated assets.

In both cases, the commercial challenge was not only creating new technology. It was establishing the trust, standards, and partnerships required for a new market to work.

When AI moves from answers to actions

The need for stronger verification is growing as AI agents move deeper into enterprise software.

Gartner predicts that up to 40 percent of enterprise applications will include task-specific AI agents by the end of 2026, compared with less than 5 percent in 2025. These systems are increasingly being designed to complete workflows, interact with software, and make operational decisions, rather than simply generate text.

According to Gartner, 40% of enterprise applications are expected to be integrated with task-specific AI agents by the end of this year, up from less than 5% a year ago. The advisory company also predicts that, in its best-case scenario,agentic AI will drive approximately 30% of enterprise application software revenue by 2035, up from a mere 2% in 2025.

"There is a major difference between an AI giving you a bad answer and an AI taking the wrong action," Koverko said. "Once an agent can approve something, change a database, contact a customer, or move money, unreliable information becomes an operational risk."

That risk is already showing up in enterprise surveys. AvePoint's 2026 report, based on responses from 750 technology, security, and AI leaders, found that 88.4 percent of organizations reported at least one AI agent-related security breach during the prior twelve months. Data leakage and manipulation through malicious or untrusted inputs were among the most common incidents.

Security controls can determine what an agent is allowed to access. Logs can show what it did. But neither necessarily establishes whether the agent's underlying judgment was correct.

"Permissions tell you whether an agent was allowed to act. Logs tell you what it did," Koverko said. "They do not prove that the decision was right. That requires a separate quality and verification layer."

The data problem is becoming a verification problem

More data still matters, but raw volume no longer guarantees better results. The most valuable inputs are increasingly private, specialized, current, and dependent on expert judgment. They cannot simply be collected from the public internet at scale.

This is especially true in fields such as medicine, law, finance, security, and robotics, where the hardest examples are often ambiguous, and the cost of a wrong answer is high.

"The scarce input is not another billion generic data points," Koverko said. "It is qualified human judgment on the cases where the model is uncertain, and the outcome matters."

Poor quality data is not merely useless. It can actively degrade a system. A wrong label, weak evaluation, or unqualified review can teach a model the wrong lesson, influence later outputs, and become difficult to trace once it moves through a training or production pipeline.

The problem is therefore not just finding more people to review AI work. Enterprises need to define what quality means, identify which reviewers are qualified, measure agreement, and preserve a record of how the conclusion was reached.

That is the gap Sapien is targeting.

Turning subjective review into verifiable evidence

Sapien describes Proof of Quality, or PoQ, as a verification layer for AI work. The customer defines quality through a rubric. Independent validators score the same work against that standard. Consensus produces a result, and the rubric, outcome, and level of agreement are preserved in an auditable record.

The system is designed to verify training data, model evaluations, AI-generated outputs, audits, and other work where quality cannot be established by a simple automated check.

"Most quality control still ends with one company saying, 'We checked it,'" Koverko said. "That may be enough for low-risk work. It is not enough when another company, auditor, regulator, or customer needs to rely on the result."

Sapien says its network includes more than 1 million contributors across more than 110 countries and has completed more than 195 million tasks. The company has publicly identified organizations including Toyota, Alibaba, Midjourney, and the United Nations among its enterprise relationships.

That network gives Sapien access to broad human participation, but Koverko says the greater opportunity lies in routing the right level of judgment to the right problem. AI can process the obvious cases, while qualified people focus on ambiguity, exceptions, and high-consequence decisions.

"The goal is not to put a person in front of every AI action," he said. "That would be too slow and too expensive. The goal is to use people where judgment creates the most value, then make the result consistent and auditable."

A focused test of Proof of Quality

In July, Sapien published results from a company-run trial using 250 labels from an automotive history platform. The test focused on identifying incorrect matches between messy vehicle listings and catalog entries.

Ten human reviewers and separate AI reviewer panels scored the labels against the same rubric. Sapien then manually checked the 32 labels that produced the most doubt. Nine were found to be incorrect.

According to Sapien, human consensus flagged all 9 incorrect labels but none of the 23 correct labels in the reviewed group. Correcting the nine errors improved one model's performance on difficult cases from 70 percent to 82 percent and another model's performance from 86 percent to 93 percent.

The trial also suggested a practical division of labor between AI and human reviewers. An AI panel used as an initial filter reduced the volume requiring human review by about 74 percent without excluding any of the errors later found by the human process.

The test was narrow and should be understood that way. It covered one product family, and one person performed the final manual check. It was not an independent validation of PoQ across every use case. But it demonstrated the commercial thesis Sapien is now taking to potential customers: use automated systems for volume, reserve human expertise for disputed or consequential cases, and preserve evidence of how the final verdict was reached.

"The value is not the abstract idea of better data," Koverko said. "The value is catching mistakes before they enter a model or workflow, reducing the amount of expensive human review, and giving the customer a record they can defend."

Selling adoption, not replacement

Koverko's near-term focus is getting companies to test PoQ on real work. The approach is deliberately narrow. A potential customer brings one important queue of disputed labels, evaluations, or AI outputs. Sapien helps translate the customer's existing quality standard into a rubric, coordinates independent review, and returns consensus results, disagreement analysis, and a Proof Report.

The aim is not to replace a company's labeling provider, evaluation platform, observability tools, or internal quality team. PoQ is intended to sit above those systems as an independent verification layer.

"Customers do not want to rebuild their entire stack to test a new idea," Koverko said. "We want to start with one real problem, show whether PoQ catches something their current process missed, and prove the return on investment. If it works, we expand from there."

That positioning reflects Koverko's broader role at Sapien. He is helping identify the right early customers and partners, sharpen the commercial offer, and communicate why verification deserves its own layer in the AI infrastructure market.

It also makes the company's thesis more specific than the broad claim that AI needs better data. Sapien argues that AI teams need a way to demonstrate why particular data, evaluations, and outputs should be trusted, especially when there is no clear ground truth.

"We are not moving away from scale," Koverko said. "We are moving into a phase where scale has to be matched by proof. The next winners in AI will not just produce more data and more outputs. They will be able to show how quality was defined, who verified the work, and why the result can be trusted."