Kyanos · Paper VII · Applied Research · August 2026 · Revision 2 · 12 August 2026
The Limits of Remediation
Why factual errors in AI platforms can be patched, why affective framing cannot, and what that means for campaigns and endorsing organizations.
Ed Forman · Founder, Raise Presence · August 2026
Abstract
Paper V established that answer engine answers function as the uncontrolled last mile of every campaign channel. This paper takes the next step: asking what can actually be done about errors once they appear. The answer divides cleanly. Factual errors — wrong dates, wrong positions, missing endorsements — have a defined remediation path: correct the underlying sources, surface authoritative content, and outputs eventually shift. Affective framing — the tone a model takes toward a candidate, the warmth or coolness of its language, the volunteered caveats — does not. It is a downstream artifact of training data composition, safety tuning, and retrieval context, none of which a campaign can directly touch. This paper maps the available levers from most to least tractable, names the irreducible residual, and argues that the residual itself is the most important strategic finding for clients: affective drift must be detected early and routed around, not corrected at the source.
Every audit of an AI platform's treatment of a candidate produces two kinds of finding, and they are not the same kind of problem. The first is factual: the model says the candidate voted against a bill they voted for, lists an endorsement that was withdrawn, or states a position the candidate has publicly reversed. The second is affective: the model is technically accurate but consistently cooler toward one candidate than another, hedges more, volunteers more unsolicited caveats, declines to answer questions it answers freely about peers, or describes a policy position in language that subtly disfavors it.
Campaigns and endorsing organizations have learned to ask about the first kind. They have not yet learned to ask about the second. This is partly because affective error is harder to see — a single response can look unobjectionable, and the asymmetry only emerges across hundreds of queries — and partly because the remediation question for affective error is genuinely harder. The first half of this paper explains why. The second half explains what to do about it anyway.
The distinction between factual and affective error is not new to the academic literature. It is, however, only recently mature enough to support an applied practice. The most direct precedent is Bang and colleagues' 2024 work proposing that political bias in large language models be measured along two axes — content (the substance of what is said) and style (the lexical polarity and framing of how it is said).1 Their content-and-style framework maps closely onto what this paper calls the factual-and-affective distinction. The methodological contribution that follows is operational rather than foundational: turning the academic split into a measurement and remediation practice that campaigns can act on inside an election cycle.
Two further strands of research bear on the practice. First, the field has converged on the use of large language models themselves as judges of affective signal in political content, with a growing literature validating LLM-extracted measures of stance, affect, and polarization against human-coded gold sets.2 Stanford's OpinionQA work demonstrates that model output can be systematically compared to public-opinion benchmarks, providing precedent for the peer-set comparisons that any serious affective audit requires.3 Second, the underlying vocabulary for describing affective presence — warmth, competence, sincerity, and related dimensions — is well established in social psychology and consumer research, with measurement scales validated over decades.4 Applied work in this paper draws on those scales as scaffolding, with political-domain operationalization built on top.
That phrase, political-domain operationalization, carries more weight than it appears to. Substantially all of the applied optimization research now circulating under the heading of answer engine optimization was measured on general web and commercial content. The most-cited work in the field evaluated roughly 10,000 queries across 25 general-web domains such as Arts, Health and Games; the nearest thing to politics among its category tags is a general "Law & Government" label, and the authors' own conclusion is that efficacy varies by domain and that domain-specific methods are therefore needed.7 Their results bear that out internally: the technique that performed best on factual queries was not the one that performed best on debate-style queries.
Where election content has been measured directly, the figures resist comparison with the general-market numbers in circulation, and that resistance is itself instructive. An audit of ChatGPT and Google AI Overviews on voter questions in Arizona, Michigan and Pennsylvania, conducted in early 2026, found Wikipedia supplying 12.3% of all links returned. The widely quoted general-market rankings of most-cited domains put Wikipedia far higher, but they report a different quantity: how often a domain appears at all, rather than what share of the links it supplies. Their published percentages sum to well over 100%, which is the tell. The two figures cannot be set against each other, and a practice that treats them as interchangeable is borrowing a number it has not read.8 The same audit found ChatGPT returning incomplete candidate lists for 88.9% of gubernatorial queries.8 Incompleteness of that kind has no counterpart in the commercial corpora the optimization literature is built on, because a product query carries no roster that must be exhaustive.
The platforms single out this domain too, and they do not do it identically. OpenAI's usage policies, effective 29 October 2025, list "political campaigning, lobbying, foreign or domestic election interference, or demobilization activities" among prohibited uses, in the same clause as fraud, impersonation, and the automation of high-stakes decisions in areas such as medical, legal and essential government services; no other promotional sector appears in that list. Anthropic's policy, effective 15 September 2025, gives democratic processes and targeted campaign activities a named section of their own, prohibiting particular practices (personalized vote targeting built from individual profiles, communications to voters at scale that conceal their artificial origin, advocacy material resting on fabricated claims) rather than campaigning as a category.9 These restrictions govern what may be done with a service rather than how that service answers a question about a candidate, and the two should not be conflated. What they establish is narrower and still material: two vendors write rules for this domain that they write for no other marketing sector, and the rules differ, so what is permitted varies by engine.
The structural differences run the same way. The subject is a named person rather than a brand, which makes entity disambiguation a first-order problem. A funded opponent is publishing contrary material about the same entity at the same time, which is uncommon in the commercial settings this research covers. The sources carrying the most weight (encyclopedic and civic reference works, filings, local press) are precisely the ones the subject is constrained from editing directly. And the work is bounded by an election date, so a lever acting on an unknown schedule may act too late to matter. Findings imported from general-market research are evidence about general markets. Whether they hold for a candidate in a contested race is an open question, and this paper treats it as one.
The architecture an applied practice requires has three distinct layers, and conflating any two of them produces measurable degradation in the work. The first is a signal layer that watches the broader media environment — news, broadcast, ad libraries, search trends, social platforms — for attack signals and emerging narratives, and predicts the voter queries those signals are likely to generate. The second is a capture layer that systematically executes those predicted queries (and a baseline battery of standing queries) against AI platforms across time and stores raw responses. The third is a scoring layer that applies domain-specific rubrics to the captured responses — factual rubrics that flag specific errors against a ground-truth record, and affective rubrics that code hedge density, refusal and deflection rates, tonal valence separated from factual valence, volunteered caveats, and cross-platform tonal deltas.
The three layers answer different questions. The signal layer asks what should we be measuring this week? The capture layer asks what is the platform actually saying? The scoring layer asks what does that mean, and how does it compare? Practices that skip the signal layer end up measuring yesterday's narrative against today's voters. Practices that fold scoring into capture end up over-engineered (everything routed through automated scoring with no human validation) or, more commonly, under-engineered (interesting anecdotes with no time-series weight). This paper assumes the three layers are built and kept distinct.
Factual errors are a leak you can sometimes patch. Affective framing is the ambient pressure of the pipe.
— The Lever Map, Section III
The available remediation levers, ordered from most to least tractable. The first three operate on the inputs the models actually consume; the next two operate on retrieval and platform-level engagement; the final entries are limit cases.
Source-side tone shaping
News coverage, Wikipedia, policy writeups, third-party profiles
Slow / Indirect
Authoritative self-description
Campaign sites, official bios, structured issue and endorsement pages
High
Endorser surface area
Warm, specific language from trusted institutions propagates into model outputs
High
Retrieval-context shaping (SEO / GEO)
Surfacing warmer authoritative sources at the top of platform retrieval
Fastest-acting
Direct platform engagement
Documented patterns escalated through platform feedback and policy channels
Limited / Real
End-user prompt engineering
Voters do not craft careful prompts; cannot be a remediation strategy
Not Available
Direct model retraining
Outside campaign control; political tuning is platform-internal
Not Available
Two observations about this map. First, the levers most under campaign control — authoritative self-description, endorser language, retrieval-context shaping — are the same activities campaigns already do for traditional press and search. The novelty is not the activity but the audience: the consumer of these signals is now a model, not only a reader. Second, every lever above the line addresses factual content well and affective framing only obliquely. None of them lets a campaign tell a model to be warmer toward its candidate.
After every lever is exercised, some affective asymmetry remains. It is shaped by training data the campaign cannot inspect, by safety tuning the campaign cannot adjust, and by retrieval and ranking signals that vary by query and platform in ways that resist systematic correction.5 This residual is not a failure of the methodology. It is a feature of the system being measured.
The Strategic Inversion
The residual is the most valuable finding the practice produces, not the least. A campaign that knows where remediation will succeed can act on the levers above. A campaign that knows where remediation will fail can route around the failure — through channel mix, message framing, surrogate amplification, or paid placement — instead of pouring resources into a fix that will not arrive in time.
The shift in framing is consequential. Most communications functions are organized around the assumption that a problem identified is a problem fixable. The contribution of this practice to the field is to insist that some problems in the AI layer are diagnostic rather than corrective, and that distinguishing the two is itself the work. A campaign treating an unfixable affective asymmetry as a fixable one is a campaign burning persuasion dollars against a wall.
↑ ContentsV. Implications for Endorsing Organizations
Endorsing organizations occupy a distinctive position in the lever map. They are simultaneously a remediation lever for the candidates they endorse — their language about a candidate becomes training and retrieval signal — and a target of affective framing in their own right. Major progressive institutions are described by answer engines in tones that vary across platforms and over time, and those descriptions affect both their institutional standing and the perceived weight of their endorsements.
This produces a shared interest that is novel in campaign politics: the endorser and the endorsed have parallel exposure to the same affective drift, on the same platforms, often in the same query sessions. A voter asking an answer engine about a Senate candidate may receive, in the same response, an affectively shaped account of the union or advocacy group that endorsed her. The reputational risk is jointly held. The remediation, where remediation is possible, is jointly conducted. The residual, where it remains, is jointly carried.
One operational consequence is worth naming. Endorsing organizations with portfolios of supported candidates have a coherence question that individual campaigns do not: are our endorsees being treated consistently across the AI layer? An affective audit run across a portfolio surfaces patterns that any single race would miss — including patterns that correlate with race, gender, geography, or ideological positioning of candidates. That portfolio-level view is, for some endorsers, the most strategically useful application of affective measurement available today.
Factual error and affective framing are different problems. They require different detection methods, different remediation strategies, and different success criteria. Conflating them produces wasted effort on both fronts.
02
Most affective framing is not directly remediable. The campaign can shape inputs to the system; it cannot shape the system itself. This is a structural limit, not a methodological one.
03
The fastest-acting lever is retrieval-context shaping. What a model says about a candidate does not hold still for a quarter: across 12 models and more than 12,000 prompts between July and November 2024, candidate–trait associations shifted over the period, in some cases tracking news events.6 That is the argument for measuring on a cadence faster than quarterly. It is not a claim that shaping retrieval context produces a change on any particular timetable — no published study establishes one, and the platforms do not document their crawl or refresh intervals.
04
Endorser language is a remediation lever, and its direction depends on how visible the subject already is. The way trusted institutions describe a candidate is part of the material these systems retrieve from. Under benchmark conditions, adding citations, quotations and statistics to a source has raised that source's visibility in generative-engine responses, and reduced it for sources already ranked first, because the benchmark's visibility measure is normalized so that all sources cited in one response sum to a fixed total.7 Visibility under those conditions is redistributive: what a lower-ranked source gains, a leading source can lose. Position of that kind is observable in elections independently of any optimization. An audit of Google results for every candidate for federal office in the 2018 cycle found incumbents' results strikingly less partisan than challengers', before anyone had attempted to shape them.10 Read that as evidence that the lever exists and that its sign is not guaranteed, rather than as a prediction of what any particular endorsement will do. Whether a given endorsement reaches a given answer is measurable per surface. Endorsements written for the model audience as well as the human audience do double duty.
05
The unremediable residual is the highest-value finding. Knowing where remediation will fail lets a campaign route around the failure rather than spend against it. Diagnostic intelligence is the deliverable, not corrective promises.
06
Endorsers and endorsed share the affective layer. The reputational risk runs in both directions across the same platforms and the same query sessions. Coordinated monitoring serves both parties more efficiently than either alone.
Paper V argued that answer engines are the last mile of every campaign channel. This paper argues that the last mile is partly pavable and partly not, and that the difference between the two is the central operational question. Measurement tells the campaign which stretches to fix, which to detour around, and which to drive carefully across, knowing the surface will not improve before election day.
This note characterizes a body of work rather than a specific study, and no single source is cited for it. The pattern described — using large language models to extract stance and affect from political discourse, validated against human-coded subsets — is common in recent computational social science, but the claim that it is "standard" is the author's characterization and is not supported here by a citation.
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). "Whose Opinions Do Language Models Reflect?" (OpinionQA). Demonstrates the use of public-opinion benchmarks as peer-set references for evaluating model output distributions.
Aaker, J. L. (1997). "Dimensions of Brand Personality." Journal of Marketing Research, 34(3), 347–356. See also Fiske, S. T., Cuddy, A. J. C., Glick, P., & Xu, J. (2002). "A Model of (Often Mixed) Stereotype Content." Journal of Personality and Social Psychology, 82(6), 878–902. Both links resolve to publisher pages and may require institutional access. The warmth-and-competence framework is widely used in social and brand perception research; its generalization to political figures and advocacy organizations is this paper's inference, not a finding of either study.
Röttger, P., Hinck, M., Hofmann, V., Hackenburg, K., Pyatkin, V., Brahman, F., & Hovy, D. "IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance" (MilaNLP, Bocconi University; published in Transactions of the Association for Computational Linguistics, 2026). Documents persistent issue-level bias patterns in state-of-the-art LLMs that are robust to standard prompting interventions, supporting the structural-residual claim made here.
Sarah Cen et al., MIT CSAIL, MIT Sloan and MIT LIDS, 2025 (lead author now at Carnegie Mellon University). Near-daily queries across 12 models on more than 12,000 prompts between July and November 2024, producing over 16 million responses; candidate–trait associations shifted over the period, in some cases tracking news events (MIT CSAIL; reported in Tech Brew, October 6, 2025). This is observational: it establishes that outputs move on sub-quarter timescales, not that any specific intervention moved them.
Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, "GEO: Generative Engine Optimization," KDD 2024; first posted November 2023. Measured on GEO-bench, the authors' own benchmark of roughly 10,000 queries across 25 general-web domains such as Arts, Health and Games, answered by a retrieval pipeline built on gpt-3.5-turbo over the top five Google results; a smaller replication on Perplexity.ai, 200 examples, appears in Appendix C.1. The benchmark contains no electoral or candidate category, and the authors note efficacy varies by domain and that domain-specific methods are therefore needed. The paper's headline figure is a visibility gain of up to roughly 40%. The result cited here is Table 2: with all sources optimized at once, adding citations raised the visibility of sources ranked fifth in search by 115.1% while reducing the top-ranked source's by 30.3% on average, because impressions are normalized so that all citations in a response sum to a constant. No published study establishes an equivalent effect on production systems for political content.
States United Democracy Center, "AI and Elections," published 4 August 2026. Two rounds of testing against ChatGPT (free tier) and Google AI Overviews: a preliminary round of 497 responses across six states, September–October 2025, and a primary round of 402 responses across Arizona, Michigan and Pennsylvania, January–February 2026. The figures cited here are from the primary round. No response in that round was coded inaccurate, a material improvement on the 2023–24 chatbot-accuracy literature and a caution against citing that literature as current. Google AI changed its output format to links-only on 2 February 2026, mid-study, after which mentions of state election sites fell to zero. Every figure of this kind is a reading taken on a date. The general-market comparison referred to here is Semrush data for June 2025, covering some 150,000 citations, as rendered by Visual Capitalist on 18 August 2025 under the column heading "citation frequency": Reddit 40.1%, Wikipedia 26.3%, YouTube 23.5%, and so on down a top-twenty list whose values sum to roughly 274%. A column summing past 100% is reporting rate of appearance, not share of citations, and the figure is widely requoted as though it were the latter.
OpenAI, "Usage Policies," effective 29 October 2025; Anthropic, "Usage Policy," effective 15 September 2025. Quoted from the published policies. This is a comparison of stated policy only: neither document describes retrieval or answer behavior, and neither restricts what a model may say about a candidate.
Danaë Metaxa, Joon Sung Park, James Landay and Jeff Hancock, "Search Media and Elections: A Longitudinal Investigation of Political Search Results in the 2018 U.S. Elections," Proceedings of the ACM on Human-Computer Interaction (CSCW '19), 2019. Roughly 4 million URLs from the first page of Google results for every candidate for federal office across six months of the 2018 cycle: more than 3,000 candidates for 225 seats, 878 of them on the general election ballot. Google only, and it pre-dates generative answer engines. The authors excluded social media, campaign websites and encyclopedia entries as unscoreable for partisanship, which is much of the surface set this paper is concerned with. Their overall finding was no partisan skew for either party, with results emphasising authoritative sources. The incumbency contrast is what is cited here; the study does not address answer engines.
About the Author
Ed Forman
Ed Forman is the founder of Raise Presence LLC, which built Kyanos to measure and improve how progressive candidates and causes are represented across AI platforms. About Raise Presence →
Note: AI helped me research and draft this document. I have read every word and it has been processed by my (human) brain — I've checked what AI drafted for me, made the edits I wanted, and I stand behind what it says.