From ranking to reasoning: when the answer starts buying
For thirty years, being found online meant one thing: a query, a ranked list, and a click. Large language models are dissolving that sequence. They answer instead of ranking, and the next generation of agents will buy instead of referring. This is a working paper: the argument, the four tests that would settle it, and what I would already be doing about it if I ran your funnel.
The click is the atom of the digital economy. It is what advertisers buy, what publishers sell, and what every analytics dashboard counts. Search advertising became the largest advertising market in the world by pricing one thing: position in a ranked list at the moment someone expressed intent. Organic search became the default acquisition channel for the same reason. Both institutions assume a user who reads the list, chooses, and clicks.
A generative engine removes the list. It reads the same web a crawler always did, but instead of returning ranked pointers it returns a synthesized answer: a paragraph that compares the options, weighs the trade-offs, and cites two or three sources where a results page exposed thousands. The interface change looks small. The economic change is not, because the answer absorbs the click. Someone who is told which tool to buy, in fluent prose with the comparison already done, has little reason to visit the eleven review sites whose work the answer summarizes.
What is actually changing?
The shift is arriving in three stages, and it helps to name them, because each stage moves a different asset and threatens a different business.
Stage one: answers inside search. A generated summary sits above the ranked list and intercepts a share of its clicks. The marginal loser is the publisher whose referral is absorbed. This stage is already here.
Stage two: answer engines. Conversation is the whole interface and the ranked list never appears. The marginal loser is the referral channel itself, and with it the attribution regime that made digital marketing accountable. If nobody clicks, nothing is tracked, and the influence a page had on a purchase becomes invisible to the tools that measure influence by clicks.
Stage three: agents that buy. The system does not recommend the standing desk. It locates it, compares merchants, fills the checkout, and completes the order under delegated authority. The marginal loser is the merchant's direct relationship with the customer, which is the asset that retention, expansion and pricing power all sit on. At stage one the model competes with the results page for attention. At stage three it competes with your storefront for the transaction.
This is the same direction of travel I wrote about in my read of Sequoia's services thesis: the machines are moving up the stack, from executing tasks to owning outcomes. Here the outcome they are learning to own is the purchase.
Why the click mattered more than it looked
Economics has always treated search as costly, and priced markets accordingly. Stigler showed in 1961 that price dispersion survives because comparing options costs effort. The information-foraging school modeled people as cost-tuned foragers who follow scent and abandon patches that stop yielding. Both traditions make the same prediction about the present: an interface that cuts search costs to almost nothing gets adopted fast, because the saving is immediate and personal while the damage is deferred and diffuse.
Here is the part that matters for business. When the cost of finding options approaches zero, the scarce thing stops being information. It becomes trust in the selector. A market where a conversational agent holds the shortlist is a market with a new gatekeeper, and the history of gatekeepers, from travel agents to app stores, is consistent: the rents move to the gate. Platform economics has formalized this since Rochet and Tirole. A search engine was a restrained gatekeeper, because the ranked list left the choice, and the customer relationship, with you. An answer engine holds the shortlist. A purchasing agent holds the transaction. That is the strongest gatekeeper position the consumer internet has produced.
What the research already says
The academic groundwork exists, in three separate rooms that rarely talk to each other. Full citations are at the bottom of the page.
The capability is real. Google's own researchers argued in 2021 that retrieval systems should behave like domain experts answering directly rather than librarians pointing at shelves. Retrieval-augmented generation made the proposal practical by grounding a model's answers in a live corpus. Conversational search then changed what a query is: a dialogue can host the whole funnel, from vague problem to comparison to choice, where a ranked list served one lookup at a time.
Visibility in answers is measurable and partly controllable. The GEO paper at KDD 2024 showed that deliberate content changes, quotable claims, cited statistics, clear structure, shift a source's inclusion in generated answers by meaningful margins, up to forty percent on their visibility metrics. Two things follow. Inclusion is a controllable variable, which makes it a strategic investment rather than a lottery. And wherever inclusion is controllable, an optimization industry forms. It already has a name, and I use its methods on this site: every article here is written to be quotable by an answer engine as much as readable by you.
Trust behaves strangely in dialogue. The advice literature holds an apparent contradiction. People abandon algorithms after seeing them err, which Dietvorst and colleagues named algorithm aversion. People also weight algorithmic advice above human advice in many judgement tasks, which Logg and colleagues named algorithm appreciation. The resolution is in the conditions, and conversational interfaces manipulate exactly those conditions: dialogue personalizes, fluent prose conceals uncertainty, and the answer arrives without the visible disagreement between sources that a results page shows by default. Every one of those mechanisms plausibly pushes confidence above accuracy. And there is a second, quieter problem: consumers have spent decades learning to recognize ads and discount them. A synthesized recommendation carries none of the markers that trigger that defense. If paid influence enters the answer layer, it may inherit the credibility of advice rather than the skepticism reserved for advertising.
The pricing theory no longer fits. Position auctions, the mechanism that made search the most efficient advertising market ever built, presume discrete slots, observable clicks, and an advertiser bidding for a user who still makes the final choice. An answer has no slots, absorbs the click, and at the agentic stage makes the choice itself. Influence over a shortlist assembled inside a model will still be priced. Without a transparent auction, it will be priced opaquely.
Four tests that would settle it
I have not run these studies. I am writing the predictions down before the evidence, because a prediction dated after the data is just a story. If someone with the data wants to run any of these with me, my inbox is open.
Test one: substitution, by intent. Take a panel of sites with server-side analytics and exploit the fact that generative answers rolled out market by market and query type by query type. Compare click-through per impression before and after, against untouched query classes as the control. My prediction: real click losses, concentrated on informational and comparison queries, smallest on navigational and transactional ones. And a second signature: the share of branded queries rises, because the click that survives the answer is a verification click. People check what the machine told them by searching for the name it gave.
Test two: who gets cited, and can you move it. Run a fixed battery of commercial questions against the major answer engines every week and log which domains get cited. Then randomize legitimate content improvements, structured data, question-shaped headings, quotable one-line claims, across matched pages and watch inclusion. My prediction: citation is far more concentrated than the results pages ever were, a handful of domains carrying most answers, and inclusion responds to structure, meaningfully but less than domain authority does. The week-to-week churn of the cited set is the number I would watch most, because it tells you whether answer visibility is an asset you accumulate or a treadmill you rent.
Test three: what conversation does to buyers. Give three groups the same purchase task with real money: one searches a ranked catalog, one asks a grounded assistant, one can delegate the checkout to it. Measure how many options people consider, the quality of what they choose, what they pay, and how confident they are versus how right they are. My prediction: conversation shrinks the shortlist and widens the gap between confidence and accuracy. Sponsorship labels that work on a results page will work less well inside a dialogue. And people who delegate will pay more than people who choose by hand from the same recommendations.
Test four: follow the margin. Sit with merchants who have plugged into agent-mediated checkout and read the actual fee structures. My prediction: agent orders carry a toll that direct orders do not, the customer record accrues to the agent's operator rather than the merchant, and liability for errors lands on the merchant while influence over discovery stays with the operator. The exceptions, merchants who kept the customer relationship, will be the most instructive cases in the whole program.
Search is becoming an answer. The answer is becoming an agent. Each step moves attention, then trust, then the transaction itself inside the machine's mediation, and each step reprices a different asset on your balance sheet.
What this means for your business
You do not need to wait for the studies to act on the direction. Four moves follow, and they happen to be the four I sell, which is either a conflict of interest or the reason I noticed. You can judge.
Treat machine legibility as a distribution channel. Content written to be cited, not just ranked: quotable single-sentence claims, statistics with named sources, questions as headings, structured data underneath, and a crawler policy decided as channel strategy rather than left as a server default. This runs parallel to classical SEO and will outgrow it. It is also cheap right now, the way SEO was cheap in 2003, and it will not stay cheap.
Rebuild brand as the retrieval prior. When the interface hides the shelf, the name the customer asks for is the last channel no shortlist can intercept. Demand that arrives pre-named cannot be disintermediated by a consideration set. If my substitution prediction is right, branded search share is about to become the honest KPI of marketing, because it measures the only demand the answer layer cannot take from you. This is positioning work, and it just became more valuable, not less.
Price as if every comparison will be made perfectly. An agent reads every pricing page in your category without fatigue and re-compares on every purchase. Packaging built on obscurity, decoy tiers, hidden fees, comparison friction, is a dated liability. Packaging built on genuine differentiation, and on relationships an agent cannot arbitrate, keeps its power. If your margin depends on the customer not looking too closely, fix that before the agents arrive, on your schedule rather than theirs.
Measure answer share, not just rankings. Referral attribution undercounts your influence in exact proportion to the substitution. The successor metric is answer share: how often, and how favorably, you appear in the generated answers to your category's questions, tracked the way consumer brands track share of shelf. The weekly prompt battery from test two is not just a research instrument. It is the prototype of a dashboard your marketing team should already have.
Where this argument breaks
Three honest weaknesses. First, engines drift: model versions change without notice, so anything you measure about answer behavior is a dated observation of a moving system, and any visibility you win can be rearranged in a release you were not consulted on. Second, the middle of my argument leans on adoption continuing, and adoption is a bet, not a law. If answer quality stalls, or a hallucinated recommendation produces a lawsuit ugly enough to slow delegation, the agentic stage arrives late or arrives regulated. Third, the answer economy has an unpriced dependency: it feeds on content that referral revenue currently funds. If stage two collapses that revenue, the corpus that grounds the answers degrades. Nobody, including the engine operators, has a good story for that yet.
None of these change the direction. They change the timetable, and timetables are what operators plan against.
Sources
The research this paper stands on, verified against the published record.
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K. and Deshpande, A. (2024). GEO: Generative engine optimization. Proceedings of the 30th ACM SIGKDD Conference. doi:10.1145/3637528.3671900
- Callaway, B. and Sant'Anna, P. H. C. (2021). Difference-in-differences with multiple time periods. Journal of Econometrics, 225(2), 200–230. doi:10.1016/j.jeconom.2020.12.001
- Dietvorst, B. J., Simmons, J. P. and Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126. doi:10.1037/xge0000033
- Edelman, B., Ostrovsky, M. and Schwarz, M. (2007). Internet advertising and the generalized second-price auction. American Economic Review, 97(1), 242–259. doi:10.1257/aer.97.1.242
- Lewis, P., Perez, E., Piktus, A. et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. proceedings.neurips.cc
- Logg, J. M., Minson, J. A. and Moore, D. A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151, 90–103. doi:10.1016/j.obhdp.2018.12.005
- Metzler, D., Tay, Y., Bahri, D. and Najork, M. (2021). Rethinking search: Making domain experts out of dilettantes. ACM SIGIR Forum, 55(1), 1–27. doi:10.1145/3476415.3476428
- Pirolli, P. and Card, S. (1999). Information foraging. Psychological Review, 106(4), 643–675. doi:10.1037/0033-295X.106.4.643
- Rochet, J.-C. and Tirole, J. (2003). Platform competition in two-sided markets. Journal of the European Economic Association, 1(4), 990–1029. doi:10.1162/154247603322493212
- Stigler, G. J. (1961). The economics of information. Journal of Political Economy, 69(3), 213–225. doi:10.1086/258464
- Varian, H. R. (2007). Position auctions. International Journal of Industrial Organization, 25(6), 1163–1178. doi:10.1016/j.ijindorg.2006.10.002
Common questions
Will LLMs replace search engines?
LLMs are replacing the interface of search rather than the plumbing: crawling and retrieval still happen underneath, but the ranked list of links is being replaced by a synthesized answer, and answers absorb most of the clicks a list used to send out. Expect substitution to be strongest on informational and comparison queries and weakest on navigational ones, where people already know where they are going.
What is generative engine optimization?
Generative engine optimization, or GEO, is the practice of making content more likely to be included and cited in AI-generated answers. Peer-reviewed work has shown that quotable single-sentence claims, cited statistics, clear structure and question-shaped headings measurably raise a page's visibility in generated answers. It runs parallel to classical SEO and optimizes for being quoted rather than merely ranked.
What is agentic commerce?
Agentic commerce is purchasing executed by an AI agent under delegated authority: the agent finds the product, compares merchants, fills the checkout and completes the order. It moves the transaction, and usually the customer record, from the merchant's storefront to the agent's operator, which is why it changes pricing power and not just interfaces.
How should a business prepare for AI-driven search?
Four moves: invest in machine legibility so answer engines can quote you, structured data and citable claims included; build brand, because demand that arrives pre-named is the one channel a machine's shortlist cannot intercept; price as if every comparison will be made perfectly, since agents read every pricing page without fatigue; and start measuring answer share, how often you appear in generated answers for your category's questions, alongside classical rankings.
Does AI-mediated buying make consumers better off?
It depends on a gap you can measure: if conversational recommendations shrink the options considered while keeping choice quality high, that is efficient advice; if they shrink options while confidence outruns accuracy, that is a persuasion channel bypassing defenses people learned on ads. Early evidence on trust in dialogue suggests the risk is real, which is why disclosure rules built for labeled ad slots need rethinking for answers.