From Web Search to AI Search: A Comparative Analysis of Information-Seeking Behaviour in Traditional Search Engines and Conversational Systems Based on Large Language Models

  • Arianna Di Vittorio Department of Economics, Management and Business Law, University of Bari “Aldo Moro”, Bari, Italy
  • Vito Alessandro Di Gioia MatiPay Srl, Bari, Italy
Keywords: AI search; information-seeking behaviour; Large Language Models; search engines; human-computer interaction; key performance indicators; New Customer Experience

Abstract

AI search is changing how users seek information: instead of typing keywords and clicking through links, users increasingly ask complex questions and obtain direct answers on the results page. This shift, together with the rise of “zero-click search” and of chatbots based on Large Language Models (LLMs), is reshaping the customer experience of information seeking. This paper examines how LLM-based search systems are changing the way users retrieve information online, moving the dominant paradigm from keyword queries and ranked lists of links to direct answers in natural language. The study offers a structured, non-systematic narrative review comparing traditional search engines (represented by Google Search) and conversational AI tools (represented by ChatGPT and GPT-3/3.5-based search tools). It synthesises eight independent studies published between 2023 and 2026: three randomised controlled experiments, two large-scale observational analyses of real-world behaviour and three studies measuring content reliability, complemented by an industry benchmark (Vectara Hallucination Leaderboard). The analysis focuses on five key performance indicators (KPIs): number of queries per task (KPI 1), task success rate/accuracy (KPI 2), total time to task completion (KPI 3), rate of consultation of external sources/outbound clicks as a proxy for cross-verification (KPI 4), and the hallucination rate of AI systems compared with the rate of unreliable results in traditional search (KPI 5). AI systems consistently reduce task completion time, whereas the reduction in the number of queries depends on the task. Accuracy improves only when the generated output is correct; when it is wrong, accuracy drops sharply, consistent with a markedly lower consultation of external sources. The risk of error exists in both paradigms under specific conditions of use. The five KPIs should therefore be read as an interdependent system, with implications for the design of hybrid search tools in human-computer interaction and information retrieval.

Downloads

Download data is not yet available.

References

Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5–16). ACM. https://doi.org/10.1145/3637528.3671900

Alter, A. L., & Oppenheimer, D. M. (2009). Uniting the tribes of fluency to form a metacognitive nation. Personality and Social Psychology Review, 13(3), 219–235. https://doi.org/10.1177/1088868309341564

Aslett, K., Sanderson, Z., Godel, W., Persily, N., Nagler, J., & Tucker, J. A. (2024). Online searches to evaluate misinformation can increase its perceived veracity. Nature, 625(7995), 548–556. https://doi.org/10.1038/s41586-023-06883-y

Borenstein, M., Hedges, L. V., Higgins, J. P. T., & Rothstein, H. R. (2009). Introduction to meta-analysis. Wiley. https://doi.org/10.1002/9780470743386

Campbell, M., McKenzie, J. E., Sowden, A., Katikireddi, S. V., Brennan, S. E., Ellis, S., Hartmann-Boyce, J., Ryan, R., Shepperd, S., Thomas, J., Welch, V., & Thomson, H. (2020). Synthesis without meta-analysis (SWiM) in systematic reviews: Reporting guideline. BMJ, 368, l6890. https://doi.org/10.1136/bmj.l6890

Caramancion, K. M. (2024). Large language models vs. search engines: Evaluating user preferences across varied information retrieval scenarios (arXiv:2401.05761). arXiv. https://arxiv.org/abs/2401.05761

Chaiken, S. (1980). Heuristic versus systematic information processing and the use of source versus message cues in persuasion. Journal of Personality and Social Psychology, 39(5), 752–766. https://doi.org/10.1037/0022-3514.39.5.752

Chapekis, A., & Lieb, A. (2025, July 22). Google users are less likely to click on links when an AI summary appears in the results. Pew Research Center. https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/

Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. Theoretical Population Biology, 9(2), 129–136. https://doi.org/10.1016/0040-5809(76)90040-X

Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J. L., Clowez, G., Boileau, P., & Ruetsch-Chelli, C. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. Journal of Medical Internet Research, 26, e53164. https://doi.org/10.2196/53164

Chinn, S. (2000). A simple method for converting an odds ratio to effect size for use in meta-analysis. Statistics in Medicine, 19(22), 3127–3131. https://doi.org/10.1002/1097-0258(20001130)19:22<3127::AID-SIM784>3.0.CO;2-M

Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1), 64–93. https://doi.org/10.1093/jla/laae003

Fishkin, R. (2024, July 1). 2024 zero-click search study: For every 1,000 US Google searches, only 374 clicks go to the open web. In the EU, it’s 360. SparkToro. https://sparktoro.com/blog/2024-zero-click-search-study-for-every-1000-us-google-searches-only-374-clicks-go-to-the-open-web-in-the-eu-its-360/

Forrester. (2024). Buyers’ journey survey, 2024. Forrester Research.

G2. (2025). 2025 buyer behavior report. G2 Research. https://research.g2.com/cmos-2025-buyer-behavior-report-research-g2

Guyatt, G. H., Oxman, A. D., Vist, G. E., Kunz, R., Falck-Ytter, Y., Alonso-Coello, P., & Schünemann, H. J. (2008). GRADE: An emerging consensus on rating quality of evidence and strength of recommendations. BMJ, 336(7650), 924–926. https://doi.org/10.1136/bmj.39489.470347.AD

Harsel, L. (2025, July 30). Google AI Mode’s early adoption and SEO impact. Semrush. https://www.semrush.com/blog/google-ai-mode-seo-impact/

Hedges, L. V., Gurevitch, J., & Curtis, P. S. (1999). The meta-analysis of response ratios in experimental ecology. Ecology, 80(4), 1150–1156. https://doi.org/10.1890/0012-9658(1999)080[1150:TMAORR]2.0.CO;2

Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), 1–55. https://doi.org/10.1145/3703155

Iannelli, M., & Ai, A. (2026). The new shape of search: How conversational AI recomposes information seeking (arXiv:2607.04282v3, version of 25 August 2026; v1 of 5 July 2026). arXiv. https://arxiv.org/abs/2607.04282

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1–38. https://doi.org/10.1145/3571730

Kaiser, C., Kaiser, J., Schallner, R., & Schneider, S. (2025a). How generative AI is transforming consumer decision-making. NIM Insights Research Magazine, 7. Nuremberg Institute for Market Decisions.

Kaiser, C., Kaiser, J., Schallner, R., & Schneider, S. (2025b). A new era of online search? A large-scale study of user behavior and personal preferences during practical search tasks with generative AI versus traditional search engines. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA ’25). ACM. https://doi.org/10.1145/3706599.3720123

Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80. https://doi.org/10.1518/hfes.46.1.50_30392

Lemon, K. N., & Verhoef, P. C. (2016). Understanding customer experience throughout the customer journey. Journal of Marketing, 80(6), 69–96. https://doi.org/10.1509/jm.15.0420

Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. In Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 7001–7025). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.467

Marchionini, G. (2006). Exploratory search: From finding to understanding. Communications of the ACM, 49(4), 41–46. https://doi.org/10.1145/1121949.1121979

Metzler, D., Tay, Y., Bahri, D., & Najork, M. (2021). Rethinking search: Making domain experts out of dilettantes. ACM SIGIR Forum, 55(1), 1–27. https://doi.org/10.1145/3476415.3476428

Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055

Pirolli, P., & Card, S. (1999). Information foraging. Psychological Review, 106(4), 643–675. https://doi.org/10.1037/0033-295X.106.4.643

Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. https://doi.org/10.1016/j.tics.2016.07.002

Shah, C., & Bender, E. M. (2022). Situating search. In Proceedings of the 2022 ACM SIGIR Conference on Human Information Interaction and Retrieval (CHIIR ’22) (pp. 221–232). ACM. https://doi.org/10.1145/3498366.3505816

Shi, Q., Zhu, K., & Gu, K. (2026). Answering without referring: How AI search rewrites the web’s economic bargain (arXiv:2607.07652v1, version of 8 July 2026). arXiv. https://arxiv.org/abs/2607.07652 (also deposited on SSRN, abstract no. 7035298, https://doi.org/10.2139/ssrn.7035298)

Spatharioti, S. E., Rothschild, D., Goldstein, D. G., & Hofman, J. M. (2025). Effects of LLM-based search on decision making: Speed, accuracy, and overreliance. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). ACM. https://doi.org/10.1145/3706598.3714082

Stigler, G. J. (1961). The economics of information. Journal of Political Economy, 69(3), 213–225. https://doi.org/10.1086/258464

TrustRadius. (2025). Bridging the trust gap: B2B tech buying in the age of AI. https://go.trustradius.com/rs/827-FOI-687/images/TrustRadius-Bridging-the-Trust-Gap-B2B-Tech-Buying-in-the-Age-of-AI.pdf

Vectara. (2026). Hallucination leaderboard (HHEM-2.3, update of 11 May 2026) [GitHub repository]. Retrieved October 3, 2026, from https://github.com/vectara/hallucination-leaderboard

Xu, R., Feng, Y., & Chen, H. (2023). ChatGPT vs. Google: A comparative study of search performance and user experience (arXiv:2307.01135). arXiv. https://arxiv.org/abs/2307.01135

Published
2026-10-10
How to Cite
Di Vittorio, A., & Di Gioia, V. A. (2026). From Web Search to AI Search: A Comparative Analysis of Information-Seeking Behaviour in Traditional Search Engines and Conversational Systems Based on Large Language Models. European Scientific Journal, ESJ, 58, 187. Retrieved from https://eujournal.org/index.php/esj/article/view/21635
Section
ESI Preprints