When does corporate AI disclosure convey information about technological resources?
Corporate disclosure about artificial intelligence (AI) has expanded rapidly, yet much of it describes general risks. This study distinguishes specific AI capability claims, sentences that describe the firm’s own AI capability and satisfy at least three of six specificity criteria, from generic and firm-specific AI risk disclosure, and examines whether these claims track technological resources: R&D intensity, AI patent portfolios, and the AI workforce share. The associations are stronger where a resource is relevant to how the industry develops or uses AI, and litigation exposure weakens the patent-disclosure association.
The main results
Each card states one finding in words; the estimates behind it (coefficient signs and p-values, within | between firms) appear with the button and in full in the Findings view.
Resources and the three forms of AI disclosure
How each technological resource at t-1 is associated with specific capability claims and with the two comparison outcomes, generic and firm-specific AI risk. Each cell shows the two designs; hover a mark for the estimate.
AI disclosure by sector and fiscal year
Share of each sector's 10-Ks containing the selected form of AI disclosure. Hover a point for the exact share and sample size.
Hypotheses and estimates
Each estimate comes from the study's result files, in the same two designs throughout: within firm (firm and fiscal-year fixed effects, comparing a firm with itself as its resources change) and between firms (sector × fiscal-year fixed effects, comparing firms within the same sector and year). Bands and whiskers are 95% intervals (estimate ± 1.96 standard errors, clustered by firm). All estimates are conditional associations; disclosure is chosen, not assigned.
H1: technological resources and specific capability claims
The association of each resource at t-1 with AI disclosure at t. Filled dot: within firm; hollow dot: between firms. The generic (G) and firm-specific (F) AI risk rows are comparison outcomes: the same models with risk language as the outcome, for which resource theory makes no prediction. R&D intensity is associated with claims rather than risk language; AI patent portfolios are associated with every form of AI disclosure.
H2: industry context and the resource associations
Interaction terms: AI patent portfolio × internal AI development (S₂), and R&D intensity × high industry AI exposure (S₁).
H3: litigation exposure, screening and chilling
The resource slope on specific capability claims as litigation exposure rises (the share of the sector's firms named in securities class actions in the calendar year before the filing). Screening (H3a) predicts a rising slope, chilling (H3b) a falling one. The AI patent slope falls in both designs; the R&D slope rises within firm only.
Resource and litigation associations by sector
The resource slope estimated separately for each sector, drawn where at least 20 firms hold the resource. Green: sectors classified as developing AI internally; amber: sectors that obtain AI externally. Hover for the estimate, p-value and firm count; the Wald test assesses whether the drawn slopes are equal.
Nine sectors, two ways of obtaining AI
In five sectors internal AI development is relatively important (software and IT services, computers and chips, aerospace and defense, auto manufacturing, and pharma and biotech); four primarily obtain AI through external providers (retail, utilities, construction, and construction machinery). This classification is the congruence moderator S₂: an AI patent portfolio is expected to be more closely associated with capability claims where AI is developed internally.
The filings behind the measures
- Each row is a company and each square one annual report (10-K), colored by the specific AI capability claims it contains.
- Click a square to open that report in a window: every AI sentence the three coders found, grouped by type with its six specificity criteria, beside the firm’s R&D, AI patents, AI workforce and litigation exposure.
- Click a company name to open its most recent report with AI sentences; inside the window, the year buttons or the ← → keys move between its reports.
- Hover a square for a one-line preview; Ctrl-click (⌘-click on a Mac) opens the 10-K itself on sec.gov.
- Search finds a company in any of the nine sectors; the menu switches sectors.
- A short tour plays when this tab opens. Click anywhere to stop it and explore, or press Play tour to watch it again.
Loading the filings grid…
How it was measured
Data
Every original 10-K filed on SEC EDGAR by US registrants in nine SIC-defined sectors, fiscal years 2014 to 2025, frozen under a manifest. Resources and controls come from Compustat and SEC XBRL at t-1; AI patents from the USPTO Artificial Intelligence Patent Dataset and PatentsView, accumulated with 15% annual depreciation; the AI-worker share from the replication package of a published study (available to fiscal year 2022); securities class actions from Audit Analytics and the Stanford Securities Class Action Clearinghouse; industry AI exposure from published occupation-based scores.
From text to variables
A lexicon of 13 AI term families (period-consistent, so the post-2022 jump is not manufactured by today's vocabulary) pulls every AI-term sentence with its neighbors. Three open-weight language models from three labs (Alibaba, Mistral AI, Microsoft) code each sentence locally: is it about AI, does it claim a capability, which of six specificity points does it meet (action, use case, named product, number, date or stage, verifiable detail), and is it generic or firm-specific AI risk? A label needs 2-of-3 agreement. A capability sentence with 3 or more points is a specific capability claim (C); risk sentences become generic (G) or firm-specific (F) AI risk. Counts are scaled per 10,000 words of filing text.
The designs
Within firm: firm and year fixed effects, so a firm is compared with itself as its resources change. Between firms: sector × year fixed effects, so a firm is compared with its sector peers in the same year. Controls at t-1: size, leverage, cash, capex, ROA, R&D indicators and 10-K length at t. Standard errors are clustered by firm. The operating screen keeps firms with at least $10M of assets or revenue and a 10-K of at least 5,000 words.
What this is not
A 10-K records what a firm chose to tell investors. Every result on this site is a fact about corporate language and the resources behind it, not a measure of site-level technology adoption, and every estimate is an association: disclosure is chosen, never assigned.
The variables, one by one
Every variable enters the models as one number per firm-year, but the raw material differs: sentences of the 10-K, lines of the financial statements, patent grants, workforce records, and securities class actions. Each panel shows the variable’s distribution over the estimation sample; hover a bar for the exact count (bar heights are square-root scaled, since most firm-years sit at or near zero). The manuscript reports the same statistics with pairwise correlations.
Apply the specificity criteria
The study's coders are three language models reading with a codebook; this panel approximates the six specificity criteria with pattern rules to illustrate what "specific" means. Paste an AI sentence from any annual report, or load an example.