What a New Carnegie Mellon-Larridin Study Reveals About AI Adoption and Revenue Growth
Every enterprise now claims an AI strategy. Almost none of them can prove it moved the business. A new working paper out of Carnegie Mellon University’s AI Capstone Program, produced in collaboration with AI-adoption analytics firm Larridin, put that gap to an empirical test across more than 500 large U.S. public companies, and the answer is more specific, and more useful to security and compliance leaders, than most AI-adoption research gets.
The paper, “Do AI Adoption Signals Predict Company Performance? Evidence from a 500-Company Cross-Section of Scores, Filings, and Hiring Data” (Li, Tao, Xu, and Hong, July 2026 working draft), tested whether AI-adoption signals, investment intensity, hiring activity, vendor scores, and disclosure language, actually predict revenue growth, margin improvement, or stock returns. The researchers built an evidence-first measurement pipeline, scored hundreds of 10-K filings and tens of thousands of job postings, and ran the numbers against real financial outcomes rather than survey sentiment.
The headline finding is not what most AI-maturity vendors want to hear: spending, hiring, and generic AI-adoption scores mostly wash out once you control for company size. The one signal that survived every statistical control the researchers threw at it was narrative concreteness, whether a company can point to named, deployed AI systems with quantified business results, as opposed to aspirational AI language. That is a finding with direct implications for how organizations building AI governance and compliance programs, Kiteworks included, should think about measuring, and proving, AI value.
This matters for a reason that goes beyond marketing. If the market is starting to separate genuine AI deployment from AI rhetoric, and if the winning signal is evidence of a specific, governed, production system, then the organizations best positioned to benefit are the ones that can already produce that evidence: who accessed what data, under what authorization, with what outcome. That is a data governance and audit-trail problem before it is an AI problem.
Key Takeaways
1. Specificity beats spending.
Among more than 500 large U.S. public companies, the AI-adoption signal most strongly associated with revenue growth was narrative concreteness, named, deployed use cases with quantified results, not total AI investment, vendor count, or headcount.
2. The concreteness effect is large and durable.
Moving from the least to the most concrete AI disclosure was associated with roughly 8 percentage points higher year-over-year revenue growth, a result that held after controlling for sector, company size, and prior growth momentum.
3. AI is currently a growth story, not a margin story.
The study found no meaningful relationship between AI-adoption signals and operating-margin improvement, and no evidence that public AI signals predicted stock returns over the following four months.
4. Generic AI-risk disclosure has become table stakes.
About 75% of the companies studied clustered at the same AI-risk-disclosure score, meaning boilerplate AI-risk language no longer distinguishes mature programs from immature ones.
5. The measurement discipline matters as much as the finding.
The researchers required every AI-adoption score to cite verifiable source evidence rather than model inference alone, the same evidence-first principle that should govern how enterprises measure their own AI programs and how they prove compliance to a regulator or auditor.
Inside the Study: 500+ Companies, Three Data Sources, One Question
The research team, Yixiao Li, Siru Tao, Xin Xu, and Hanzhe Hong at Carnegie Mellon, working with Larridin, built their study universe from Larridin’s AI Transformation Tracker coverage set combined with the S&P 500, arriving at 564 companies spanning all twelve sectors. To keep a handful of AI-semiconductor mega-caps from dominating the results, the primary analysis excludes the five largest AI-semiconductor firms (Nvidia, Broadcom, AMD, Micron, and Intel); the paper reports that every conclusion holds in the full sample as well.
Three independent data sources feed the analysis. Larridin’s AI Transformation Tracker assigns each company evidence-backed 1-to-5 scores across adoption, workforce proficiency, and realized impact, plus a composite maturity index, covering 562 companies. A large-language-model extraction pipeline separately scored 478 companies’ most recent 10-K filings on three dimensions: AI investment intensity, narrative concreteness, and risk-disclosure depth, with a rule that every non-null score had to cite a verbatim supporting passage from the filing. And a third team classified 30,861 individual job postings across 536 companies into AI-builder, AI-user, AI-leader, and non-AI roles, to construct an AI-hiring intensity measure.
Against those three signal families, the researchers tested three questions: do AI-adoption signals predict revenue growth, margin change, and productivity; do they predict market-adjusted stock returns over the following four months; and where in the AI value chain do any effects concentrate. Every regression controlled for sector, company size, and each firm’s pre-existing revenue-growth momentum, which addresses the obvious confound that fast-growing companies both adopt AI faster and keep growing regardless of AI.
You Trust Your Organization is Secure. But Can You Verify It?
The Signal That Held Up: Narrative Concreteness
Here is where the paper earns its relevance to anyone building an AI governance program. Every score- and disclosure-based signal in the study was associated with revenue growth in a simple, unconditional comparison. But as the researchers tightened the statistical controls, adding sector, then size, then growth momentum, most of those signals lost their statistical footing. Composite AI-adoption scores, which aggregate several components that naturally scale with company size, attenuated once size entered the model. AI investment intensity was positive throughout but lost significance once size was controlled for (p = 0.14). The AI-hiring builder rate showed no meaningful relationship with any outcome at all, and the authors note that hiring data was collected after the outcome quarter, which means it cannot be treated as a genuine predictive test in the first place.
Narrative concreteness was the exception. Its full-specification coefficient of 0.080 (p = 0.009) held after every control was applied, and it passed a Benjamini-Hochberg false-discovery-rate correction within the study’s pre-specified primary hypothesis family, a standard statistical safeguard against false positives when testing multiple signals at once. In practical terms, moving a company from the bottom to the top of the concreteness distribution was associated with an 8.0-percentage-point difference in year-over-year revenue growth, holding sector, size, and prior momentum fixed.
The researchers’ own interpretation is worth quoting directly, because it is the thesis the entire paper builds toward: “specificity is the signal.” What distinguished faster-growing firms, within the same sector and size class, was verifiable substance in their own regulatory disclosures, named systems in production, quantified results, rather than surface-level AI keyword exposure. Risk-disclosure language, now standardized across most companies, discriminated nothing. Each step toward evidence-anchored specificity added information.
What Didn’t Predict Performance: Spend, Headcount, and Hiring
It is worth sitting with what the study did not find, because the negative results are as strategically important as the positive one. AI investment intensity, essentially, how much a company says it is spending to build or buy AI, tracked with growth in raw comparisons but disappeared once the researchers accounted for company size. Larger companies simply invest more in everything, AI included, so investment volume alone was not distinguishing genuine AI performers from companies that are big and therefore spend a lot.
The AI-hiring intensity measure, built from a validated classification of more than 30,000 job postings, showed essentially no relationship with the tested outcomes (coefficient 0.026 in the univariate case, and negative once full controls were applied). The authors are candid about a design limitation here: because hiring data can only be collected in real time and cannot be reconstructed retroactively, their hiring snapshots were gathered after the financial quarter they were being compared against. That means the hiring result should be read as descriptive, not predictive.
For organizations evaluating their own AI programs, and for vendors selling AI-maturity scorecards built on spend, vendor counts, or headcount, this is a useful corrective. Those measures describe activity. They do not, on this evidence, describe business value.
AI Is, for Now, a Growth Story Rather Than a Margin Story
The study’s second major finding concerns where AI’s economic effect shows up, and where it does not. None of the tested signals, composite or disclosure-based, predicted four-month forward stock returns after the researchers corrected for testing multiple signals at once. The authors interpret this through ordinary market-efficiency logic: because every signal in the study is built from public information, markets appear to have already priced in what is knowable about a company’s AI posture from public sources. If anything, the point estimate on the composite AI score was mildly negative over the study window, consistent with a partial unwind of AI-theme stock valuations that had run up earlier in 2026.
Margin effects were similarly absent, no signal showed a meaningful association with operating-margin change. Combined with the revenue-growth results, the paper’s interpretation is that AI adoption is currently operating through a top-line growth channel rather than a near-term cost-reduction channel. That lines up with what security and compliance leaders are hearing internally: AI initiatives are being justified on new capability, faster product cycles, and better customer experience, not (yet) on headcount reduction.
One market-level result stands out as a separate, notable data point: the paper’s AI value-chain classification found that AI-infrastructure and hardware suppliers outperformed their sector- and size-matched peers by approximately 32 percentage points over the first half of 2026, a striking decomposition of that period’s AI-related stock rally, and a reminder that the “AI trade” and the “AI-adoption-predicts-fundamentals” question are two different phenomena that happened to occur in the same market at the same time.
Where the Signal Mattered Most: Physical-Asset-Heavy, Late-Adopter Companies
The paper’s AI value-chain decomposition splits the S&P 500 into four categories: AI infrastructure and hardware (36 companies), AI software and platforms (41), data-rich adopters (138), and physical-asset-heavy or late-adopter companies (285), the largest group by far. The narrative-concreteness-to-revenue-growth relationship was statistically significant specifically within that physical-asset-heavy, late-adopter group (n = 214, p = 0.025).
That is, on its face, an unremarkable statistical footnote. It is not. The authors argue, and this is the connective thread to security and governance, that this is precisely the segment where it is hardest to tell genuine AI deployment from AI rhetoric. A software company’s AI claims are relatively easy to verify against its product. A manufacturer, retailer, or industrial company’s AI claims are not. That is exactly where concrete, evidence-backed disclosure did the most work separating real adopters from companies using AI as a narrative device. For the large majority of the enterprise market that sits in this category, companies that are not AI infrastructure vendors or AI-native software platforms, the ability to demonstrate governed, verifiable AI deployment may be the single most differentiating asset available to them.
The Governance Implication: Evidence-First Measurement as Operating Discipline
The methodological choice underlying the entire study is itself the most transferable lesson for enterprise AI governance programs. The researchers built what they call an evidence-first LLM extraction protocol: every non-null AI-disclosure score had to cite a verbatim supporting passage from the source filing, and dimensions without supporting evidence were returned as null rather than guessed. In an internal audit, 87 to 90 percent of the model’s evidence citations verified as exact matches against the source text.
That is worth restating in plain terms: the researchers refused to let a model score a company’s AI maturity on inference alone. Every claim had to trace to a specific, citable piece of evidence. And the signal built most deeply on that principle, narrative concreteness, was also the one signal that survived every statistical challenge in the paper.
That is the same discipline that should govern how an enterprise measures, and proves, its own AI program, and it is where the paper’s findings stop being an academic curiosity and start being a blueprint. An internal AI data governance program built on the same evidence-first logic would be able to answer, for any AI use case: what production workflow is AI actually changing, who is using it, what data does it touch, what specific business outcome improved, can that result be quantified, and can the organization produce evidence, not an assertion, that the system is deployed and governed. An AI program that can answer those questions with citable evidence looks, by this study’s own logic, much more like the companies associated with real revenue growth than one built on AI announcements or pilot counts.
What This Means for Data Security, Compliance, Sovereignty, and AI Data Governance
The CMU-Larridin paper is careful to state what it does not show: it does not directly test whether stronger data security, regulatory compliance, data sovereignty, or AI governance causes better financial performance. That is an honest and important caveat, and any conclusion connecting governance maturity to the paper’s revenue-growth finding should be read as a strategic inference, not a causal claim the researchers established.
With that caveat stated plainly, the inference is still a useful one for CISOs and compliance leaders evaluating where AI governance investment belongs. Once AI moves from pilot to production, it begins touching real enterprise data, employees, customers, and workflows. That converts the operative security question from “do we allow AI” to a much more specific one: which AI system or agent can access which data, for which purpose, under which policy, and what can it do with what it retrieves. That is a data governance and access control problem before it is anything else, and it maps directly onto the study’s own finding that verifiable, deployment-level specificity is what actually distinguishes mature programs from immature ones.
The compliance implication follows the same logic. The study found that risk-disclosure depth, essentially, whether a company has AI-risk language in its filings at all, had almost no power to distinguish companies, because roughly 75% of the sample clustered at the same score. AI-risk boilerplate has become standardized to the point of being uninformative. That is a signal compliance teams should take seriously: an AI policy statement, a governance committee, or generic risk language in a filing is rapidly becoming table stakes rather than evidence of maturity. What should replace it is control evidence: the specific policy that applies to an AI workflow, the systems and data governed by it, the identity that initiated each action, the enforcement decision that was made, and the audit trail that lets the organization reconstruct the transaction later. For any organization selling or building compliance technology, the value proposition this study points toward is “we can prove the control operated,” not “we support the regulation.”
On sovereignty, the paper is explicit that it contains no empirical test of data residency, localization, or jurisdictional control as predictors of performance, any sovereignty conclusion here is an inference, not a finding. But the inference tracks the same direction as everything else in the study: as AI deployment becomes more concrete and operational, which is precisely the trend the paper documents, organizations increasingly need to know not just what data an AI system uses, but where that data resides, where it is processed, which provider can access it, and whether the organization can technically enforce the answer. For regulated, multinational, healthcare, defense, and financial-services organizations, data sovereignty compliance becomes more operational, not less, as AI moves from experimentation to production, it is not a separate track that runs alongside AI adoption.
Building an Evidence Trail Regulators and Auditors Will Accept
Put the paper’s three findings together, specificity beats spending, generic risk disclosure is now table stakes, and the concreteness effect is strongest exactly where AI claims are hardest to verify, and a fairly direct operating model falls out for AI governance teams. Governance maturity should be measured by real behavior, not by governance artifacts. An AI committee, an acceptable-use policy, an inventory spreadsheet, and annual training may all be necessary. None of them, on their own, prove that AI behavior is actually controlled.
A model built on the study’s own evidence-first philosophy would track, for every AI or agent-initiated action: the identity that initiated it, the specific data accessed or modified, the business purpose it served, the policy or approval that authorized it, the action actually taken, where the output or data went, the jurisdiction in which it was processed, and whether the organization can reconstruct that entire chain later on demand. That is not a theoretical wish list, it is the same standard of evidence a regulator, an assessor, or opposing counsel already expects when they ask an organization to prove a specific data access was authorized, encrypted, and logged.
This is the connective tissue between the CMU-Larridin study and how Kiteworks approaches Compliant AI and zero trust generative AI governance: enforcing attribute-based access control and policy decisions at the point where an AI system or a person touches sensitive data, and generating a unified audit trail that can answer exactly the questions the study’s evidence-first protocol demanded of the companies it scored. A Data Policy Engine that enforces those policies before data reaches an AI system or workflow, paired with a CISO Dashboard that surfaces what actually happened, converts governance from a set of documents into the kind of verifiable, production-level evidence the paper found to be economically meaningful. When AI-builder and AI-user activity runs through an integration point such as a Secure MCP Server, that same evidence trail extends to agent-initiated actions alongside human ones, which matters because the study’s own governance framework explicitly calls for tracking “human/agent action” as a single accountable category, not two separate ones.
None of this changes what the researchers found. It simply means that data compliance and governance infrastructure are not a tax on AI adoption, they are what makes the kind of verifiable, production-level deployment the study rewards sustainable once regulators, customers, and auditors start asking to see the evidence rather than the announcement.
If there is one action item in all of this, it is not “write a better AI policy” or “score higher on a maturity index.” Security and compliance teams must focus on the data layer, who or what touched it, under what authorization, and whether that can be proven later. Everything else in the study’s findings, and in this analysis, follows from that one point.
To learn more about building an evidence-first foundation for AI governance and compliance, schedule a custom demo today.
Frequently Asked Questions
It is a working paper titled “Do AI Adoption Signals Predict Company Performance?”, published in July 2026 by researchers Yixiao Li, Siru Tao, Xin Xu, and Hanzhe Hong through Carnegie Mellon University’s AI Capstone Program, in collaboration with AI-adoption analytics firm Larridin. It studies roughly 500 large U.S. public companies to test whether observable AI-adoption signals, investment, hiring, disclosure language, and vendor scores, are associated with subsequent revenue growth, margin change, and stock returns. As a working draft, it has not yet completed peer review, and its findings should be treated as an early, rigorously constructed data point rather than a settled conclusion for AI data governance and data compliance planning.
No. The study finds a statistical association between narrative concreteness in AI disclosures and subsequent revenue growth that survives sector, size, and momentum controls, but association is not causation, and the authors do not claim otherwise. They are also explicit that the paper does not test whether stronger data security, compliance, or AI data governance causes better financial performance, any conclusion connecting governance maturity to the study’s growth finding is a strategic inference, not a result the researchers established directly.
Not entirely, but they should stop being treated as primary evidence of AI maturity. The study found that AI investment intensity lost statistical significance once company size was controlled for, and AI-hiring intensity showed no meaningful relationship with the outcomes tested at all. Spend and headcount describe activity. A stronger internal scorecard, consistent with the study’s findings, should weight deployed use cases, quantified business outcomes, and audit trail evidence of governance far more heavily than budget or headcount figures.
Because its strongest finding, that verifiable, evidence-backed specificity is what separates real AI deployment from AI rhetoric, is fundamentally a data governance problem. Once AI moves into production, controlling which systems and identities can access which data, under what authorization, and with what auditable outcome becomes the practical mechanism for generating the kind of evidence the study associates with stronger performance. That is squarely a data governance and access control function, not a marketing or product-messaging one.
That generic AI-risk language, policy documents, and governance-committee announcements are rapidly losing their value as evidence of program maturity, the study found roughly 75% of companies clustered at the same disclosure score, meaning that language no longer discriminates mature programs from immature ones. What should replace it is control evidence: being able to show the specific policy that applied to an AI workflow, the identity that took the action, the enforcement decision, and a reconstructable audit trail of the transaction. That is the standard a regulator or assessor will expect, and it is the same evidence-first principle the study’s own researchers applied to their measurement methodology.
Additional Resources
- Blog Post
Zero‑Trust Strategies for Affordable AI Privacy Protection - Blog Post
How 77% of Organizations Are Failing at AI Data Security - eBook
AI Governance Gap: Why 91% of Small Companies Are Playing Russian Roulette with Data Security in 2025 - Blog Post
There’s No “–dangerously-skip-permissions” for Your Data - Blog Post
Regulators Are Done Asking Whether You Have an AI Policy. They Want Proof It Works.