From Web Search to AI Search: A Comparative Analysis of Information-Seeking Behaviour in Traditional Search Engines and Conversational Systems Based on Large Language Models
Abstract
AI search is changing how users seek information: instead of typing keywords and clicking through links, users increasingly ask complex questions and obtain direct answers on the results page. This shift, together with the rise of “zero-click search” and of chatbots based on Large Language Models (LLMs), is reshaping the customer experience of information seeking. This paper examines how LLM-based search systems are changing the way users retrieve information online, moving the dominant paradigm from keyword queries and ranked lists of links to direct answers in natural language. The study offers a structured, non-systematic narrative review comparing traditional search engines (represented by Google Search) and conversational AI tools (represented by ChatGPT and GPT-3/3.5-based search tools). It synthesises eight independent studies published between 2023 and 2026: three randomised controlled experiments, two large-scale observational analyses of real-world behaviour and three studies measuring content reliability, complemented by an industry benchmark (Vectara Hallucination Leaderboard). The analysis focuses on five key performance indicators (KPIs): number of queries per task (KPI 1), task success rate/accuracy (KPI 2), total time to task completion (KPI 3), rate of consultation of external sources/outbound clicks as a proxy for cross-verification (KPI 4), and the hallucination rate of AI systems compared with the rate of unreliable results in traditional search (KPI 5). AI systems consistently reduce task completion time, whereas the reduction in the number of queries depends on the task. Accuracy improves only when the generated output is correct; when it is wrong, accuracy drops sharply, consistent with a markedly lower consultation of external sources. The risk of error exists in both paradigms under specific conditions of use. The five KPIs should therefore be read as an interdependent system, with implications for the design of hybrid search tools in human-computer interaction and information retrieval.
Downloads
References
Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5–16). ACM. https://doi.org/10.1145/3637528.3671900
Alter, A. L., & Oppenheimer, D. M. (2009). Uniting the tribes of fluency to form a metacognitive nation. Personality and Social Psychology Review, 13(3), 219–235. https://doi.org/10.1177/1088868309341564
Aslett, K., Sanderson, Z., Godel, W., Persily, N., Nagler, J., & Tucker, J. A. (2024). Online searches to evaluate misinformation can increase its perceived veracity. Nature, 625(7995), 548–556. https://doi.org/10.1038/s41586-023-06883-y
Borenstein, M., Hedges, L. V., Higgins, J. P. T., & Rothstein, H. R. (2009). Introduction to meta-analysis. Wiley. https://doi.org/10.1002/9780470743386
Campbell, M., McKenzie, J. E., Sowden, A., Katikireddi, S. V., Brennan, S. E., Ellis, S., Hartmann-Boyce, J., Ryan, R., Shepperd, S., Thomas, J., Welch, V., & Thomson, H. (2020). Synthesis without meta-analysis (SWiM) in systematic reviews: Reporting guideline. BMJ, 368, l6890. https://doi.org/10.1136/bmj.l6890
Caramancion, K. M. (2024). Large language models vs. search engines: Evaluating user preferences across varied information retrieval scenarios (arXiv:2401.05761). arXiv. https://arxiv.org/abs/2401.05761
Chaiken, S. (1980). Heuristic versus systematic information processing and the use of source versus message cues in persuasion. Journal of Personality and Social Psychology, 39(5), 752–766. https://doi.org/10.1037/0022-3514.39.5.752
Chapekis, A., & Lieb, A. (2025, July 22). Google users are less likely to click on links when an AI summary appears in the results. Pew Research Center. https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/
Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. Theoretical Population Biology, 9(2), 129–136. https://doi.org/10.1016/0040-5809(76)90040-X
Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J. L., Clowez, G., Boileau, P., & Ruetsch-Chelli, C. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. Journal of Medical Internet Research, 26, e53164. https://doi.org/10.2196/53164
Chinn, S. (2000). A simple method for converting an odds ratio to effect size for use in meta-analysis. Statistics in Medicine, 19(22), 3127–3131. https://doi.org/10.1002/1097-0258(20001130)19:22<3127::AID-SIM784>3.0.CO;2-M
Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1), 64–93. https://doi.org/10.1093/jla/laae003
Fishkin, R. (2024, July 1). 2024 zero-click search study: For every 1,000 US Google searches, only 374 clicks go to the open web. In the EU, it’s 360. SparkToro. https://sparktoro.com/blog/2024-zero-click-search-study-for-every-1000-us-google-searches-only-374-clicks-go-to-the-open-web-in-the-eu-its-360/
Forrester. (2024). Buyers’ journey survey, 2024. Forrester Research.
G2. (2025). 2025 buyer behavior report. G2 Research. https://research.g2.com/cmos-2025-buyer-behavior-report-research-g2
Guyatt, G. H., Oxman, A. D., Vist, G. E., Kunz, R., Falck-Ytter, Y., Alonso-Coello, P., & Schünemann, H. J. (2008). GRADE: An emerging consensus on rating quality of evidence and strength of recommendations. BMJ, 336(7650), 924–926. https://doi.org/10.1136/bmj.39489.470347.AD
Harsel, L. (2025, July 30). Google AI Mode’s early adoption and SEO impact. Semrush. https://www.semrush.com/blog/google-ai-mode-seo-impact/
Hedges, L. V., Gurevitch, J., & Curtis, P. S. (1999). The meta-analysis of response ratios in experimental ecology. Ecology, 80(4), 1150–1156. https://doi.org/10.1890/0012-9658(1999)080[1150:TMAORR]2.0.CO;2
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), 1–55. https://doi.org/10.1145/3703155
Iannelli, M., & Ai, A. (2026). The new shape of search: How conversational AI recomposes information seeking (arXiv:2607.04282v3, version of 25 August 2026; v1 of 5 July 2026). arXiv. https://arxiv.org/abs/2607.04282
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1–38. https://doi.org/10.1145/3571730
Kaiser, C., Kaiser, J., Schallner, R., & Schneider, S. (2025a). How generative AI is transforming consumer decision-making. NIM Insights Research Magazine, 7. Nuremberg Institute for Market Decisions.
Kaiser, C., Kaiser, J., Schallner, R., & Schneider, S. (2025b). A new era of online search? A large-scale study of user behavior and personal preferences during practical search tasks with generative AI versus traditional search engines. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’25). ACM. https://doi.org/10.1145/3706599.3720123
Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80. https://doi.org/10.1518/hfes.46.1.50_30392
Lemon, K. N., & Verhoef, P. C. (2016). Understanding customer experience throughout the customer journey. Journal of Marketing, 80(6), 69–96. https://doi.org/10.1509/jm.15.0420
Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. In Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 7001–7025). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.467
Marchionini, G. (2006). Exploratory search: From finding to understanding. Communications of the ACM, 49(4), 41–46. https://doi.org/10.1145/1121949.1121979
Metzler, D., Tay, Y., Bahri, D., & Najork, M. (2021). Rethinking search: Making domain experts out of dilettantes. ACM SIGIR Forum, 55(1), 1–27. https://doi.org/10.1145/3476415.3476428
Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055
Pirolli, P., & Card, S. (1999). Information foraging. Psychological Review, 106(4), 643–675. https://doi.org/10.1037/0033-295X.106.4.643
Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. https://doi.org/10.1016/j.tics.2016.07.002
Shah, C., & Bender, E. M. (2022). Situating search. In Proceedings of the 2022 ACM SIGIR Conference on Human Information Interaction and Retrieval (CHIIR ’22) (pp. 221–232). ACM. https://doi.org/10.1145/3498366.3505816
Shi, Q., Zhu, K., & Gu, K. (2026). Answering without referring: How AI search rewrites the web’s economic bargain (arXiv:2607.07652v1, version of 8 July 2026). arXiv. https://arxiv.org/abs/2607.07652 (also deposited on SSRN, abstract no. 7035298, https://doi.org/10.2139/ssrn.7035298)
Spatharioti, S. E., Rothschild, D., Goldstein, D. G., & Hofman, J. M. (2025). Effects of LLM-based search on decision making: Speed, accuracy, and overreliance. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). ACM. https://doi.org/10.1145/3706598.3714082
Stigler, G. J. (1961). The economics of information. Journal of Political Economy, 69(3), 213–225. https://doi.org/10.1086/258464
TrustRadius. (2025). Bridging the trust gap: B2B tech buying in the age of AI. https://go.trustradius.com/rs/827-FOI-687/images/TrustRadius-Bridging-the-Trust-Gap-B2B-Tech-Buying-in-the-Age-of-AI.pdf
Vectara. (2026). Hallucination leaderboard (HHEM-2.3, update of 11 May 2026) [GitHub repository]. Retrieved October 3, 2026, from https://github.com/vectara/hallucination-leaderboard
Xu, R., Feng, Y., & Chen, H. (2023). ChatGPT vs. Google: A comparative study of search performance and user experience (arXiv:2307.01135). arXiv. https://arxiv.org/abs/2307.01135
Copyright (c) 2026 Arianna Di Vittorio, Vito Alessandro Di Gioia

This work is licensed under a Creative Commons Attribution 4.0 International License.


