From Discovery to Business Impact Assessment: A User-Centered Method for Analyzing Generative AI

discovery method

Introduction

The emergence of generative artificial intelligence, and more specifically chatbots, marks a turning point in the recent history of digital technology. Unlike many professional innovations that were adopted gradually, these tools have spread particularly rapidly among the general public. Their ability to generate content, search for information, assist with writing, and support problem-solving has quickly led many users to integrate them into their daily activities.

This widespread adoption now extends beyond the personal sphere. In many professional settings, employees spontaneously use generative AI tools to search for information, draft documents, prepare analyses, or automate certain tasks. Organizations also see significant potential in these tools for improving productivity, transforming business processes, and developing new services [1,2,3].

Yet a paradox emerges. While individual adoption of these technologies is widespread, integrating them at the organizational level often remains challenging. Many companies struggle to transform scattered individual uses into collective practices that create value. The difficulties encountered relate less to the intrinsic performance of the models than to their adoption in real-world work situations. The expected benefits do not always materialize, and many projects struggle to move beyond the experimental stage [4].

This situation can be explained in particular by the fact that the introduction of AI in a company is not solely a technological issue. It involves transformations in business activities, professional roles, collaboration methods, and managerial practices. End users, although directly affected by these changes, are still often not sufficiently involved in the phases of needs assessment, design, or solution evaluation [4,5,6].

Numerous studies show that the adoption of AI in the workplace depends on a set of interrelated factors. The work of Yann Ferguson [1,2], conducted primarily in industrial settings, highlights the importance of organizational, psychological, social, technical, and educational dimensions. Issues such as autonomy, recognition, trust, skill development, and the meaning of work strongly influence the acceptability of AI systems, sometimes even more so than their technical performance.

Given this complexity, LaborIA recommends starting with real-world work and involving users throughout the process of designing and deploying solutions [1]. These principles align directly with those of the Living Lab, a participatory innovation approach based on co-creation, real-world experimentation, collaboration among stakeholders, and continuous improvement of solutions [7].

In this context, evaluating AI solely based on its technical performance appears insufficient. Understanding its acceptability and adoption requires a multidimensional approach capable of simultaneously analyzing usage patterns, user perceptions, human–AI interactions, and impacts on work activities.

This article therefore proposes an operational method for evaluating generative AI inspired by the principles of the Living Lab. User-centered, this method combines needs analysis, usage observation, the study of human–AI interactions, and the assessment of business impacts. Its objective is to support the design, testing, and deployment of AI solutions tailored to the realities of the workplace and the expectations of organizations.

From Understanding Needs to Measuring Impact: An Approach Focused on Real-World Use

The Discovery Phase

Generative AI projects often begin with an analysis of the technology’s capabilities: Which tasks should be automated? Which model should be used? Which features should be developed? However, this technology-centric approach has a major limitation. It frequently leads to the design of solutions that are technically advanced but rarely used in practice.

There are many examples that illustrate this situation. Chatbots capable of correctly answering the majority of queries may be ignored by users if they do not fit into their work routines. Conversely, technically modest features can generate significant value when they address a clearly identified daily pain point. In both cases, the gap stems less from the system’s performance than from an insufficient understanding of the usage context.

Traditional approaches to requirements gathering also have certain limitations. Declarative surveys help identify users’ stated expectations but struggle to reveal actual practices. Process analyses highlight workflows but often overlook the human, social, or organizational factors that influence adoption. Finally, evaluations focused on model performance provide insights into AI capabilities without allowing us to anticipate its acceptability or its impact on business operations.

Our approach is therefore based on a different principle: starting with the actual work before turning our attention to the technology. The goal of the Discovery phase is not only to identify potential use cases, but to understand the professional situations in which AI could create value while also being effectively adopted by its users.

This phase aims to answer several questions simultaneously:

• Which use cases offer the best balance between business value, feasibility, and acceptability?

• What are the objectives and constraints of the various business lines involved?

• Which activities currently require significant time or effort?

• What are the daily pain points?

• What factors could promote or hinder the adoption of an AI solution?

To answer these questions, the method combines multiple sources of information. Individual and group interviews are used to gather users’ perceptions, expectations, and concerns. These discussions are supplemented by field observations to analyze real-world work situations and the operational constraints within which future applications will need to operate.

This immersion is particularly important in the context of generative AI. A solution may seem relevant in a demonstration environment but prove difficult to use in real-world business conditions. Mobility constraints, connectivity issues, noisy environments, regulatory requirements, and existing work habits directly influence the possible uses and acceptability of solutions.

Beyond identifying needs, the Discovery phase thus serves as an initial mechanism for mitigating project risk. It enables the prioritization of the most promising use cases, the anticipation of key barriers to adoption, and the steering of development toward solutions likely to generate sustainable value for both users and the organization.

Testing in a Controlled Environment

Once the requirements and use cases have been identified, a second challenge arises. An AI system may deliver excellent technical performance while still resulting in a poor user experience. Conversely, certain system limitations can be easily compensated for by users when they are well understood and properly integrated into their work practices.

Traditional evaluation approaches often rely on technical metrics such as success rates, response accuracy, or functional coverage. While necessary, these metrics provide little insight into the actual mechanisms of interaction between users and the AI. Two systems with comparable performance can thus yield very different adoption rates depending on the quality of the interactions they facilitate.

The goal of this phase is therefore to observe the AI in a real-world usage scenario, within a controlled environment that allows for a detailed analysis of human–AI interactions prior to any large-scale deployment. This step serves as a bridge between understanding user needs and evaluating the system under real-world conditions.

At this stage, the focus is primarily on ergonomic and interactional aspects. The challenge is no longer merely to verify whether the AI provides a correct response, but to understand how users construct their requests, interpret the responses received, and adapt their behavior as the conversation progresses.

Each interaction is thus analyzed as a complete conversational sequence. This approach allows us to study not only the final result obtained, but also the path taken to reach it. It highlights the mechanisms that promote successful interactions as well as the friction points that could degrade the user experience.

The analysis focuses in particular on the wording of queries, the vocabulary used, domain-specific implications, abbreviations, and the level of precision in requests. These observations help identify discrepancies between users’ natural language and the system’s actual comprehension capabilities.

The study of conversations also examines their dynamics. The number of iterations required to achieve the expected result, rephrasing, requests for clarification, or corrections made by the user are all indicators used to assess the fluidity of interactions. A technically correct response may indeed require multiple adjustments on the user’s part and generate a significant cognitive load.

The analysis of failure scenarios plays a central role in this phase. Unlike traditional evaluations, which are often limited to measuring an error rate, our approach seeks to understand the root causes of these difficulties. The causes may be related to the model itself, the formulation of queries, business ambiguities, mutual misunderstandings, or poor dialogue management. Understanding these mechanisms allows us to identify areas for improvement that are often invisible in traditional performance metrics.

This phase also allows for the creation of a corpus of conversations representative of the observed usage patterns. Beyond its value for the technical improvement of the system, this corpus provides valuable insight into the interaction strategies developed by users and the most frequently encountered situations.

Data collection is based on paired observation involving a business user and an expert observer. The user performs tasks related to their work while the expert analyzes interactions, identifies friction points, and collects immediate feedback. This setup allows for cross-referencing the objective records of conversations with the user’s subjective perception.

This phase thus provides a level of analysis often missing from traditional evaluation approaches. Whereas technical benchmarks measure the system’s capabilities and user surveys gather general perceptions, experimentation in a controlled environment allows for direct observation of the interaction between the user and the AI. It provides the necessary insights to understand not only whether the system works, but above all how it is used and why certain interactions do or do not generate value.

Field testing phase

AI can achieve excellent results in a controlled experiment without necessarily delivering the expected benefits once deployed. Real-world conditions introduce numerous variables that are difficult to replicate in a laboratory setting: organizational constraints, work habits, operational workload, interactions with other tools, and team dynamics.

The objective of this phase is therefore to evaluate the AI in its real-world environment to understand how it integrates into professional practices and what actual changes it brings to the business. It is no longer just a matter of observing human–AI interactions, but of analyzing their impact on work, the organization, and performance.

Unlike approaches focused exclusively on technical indicators or satisfaction feedback, this evaluation is based on a multidimensional analysis. It takes into account ergonomic, social, educational, organizational, ethical, and economic dimensions to understand all the factors likely to influence the adoption of the solution.

Integrating AI into daily work routines makes it possible to identify persistent barriers, operational pain points, user-driven adoption strategies, and the conditions that foster sustainable use. This insight provides a more nuanced understanding of adoption mechanisms than evaluations conducted in controlled environments.

Particular attention is also paid to the solution’s business impact. The analysis relies on indicators such as the time required to complete tasks, the quality and reliability of the results produced, the level of user autonomy, and the reduction in errors. These elements help to objectively assess the operational benefits associated with the use of AI.

Data collection relies primarily on questionnaires administered before and after the introduction of AI, combining closed-ended and open-ended questions. This comparative approach makes it possible to measure changes in user practices, performance, and perceptions throughout the deployment. It is supplemented by an analysis of the solution’s actual usage as well as interactions between users and AI, through an examination of the questions asked and the responses provided. This cross-referencing of self-reported and behavioral data provides a more nuanced understanding of adoption mechanisms, challenges encountered, and the value actually created by the solution.

Comparing “before” and “after” scenarios also facilitates the translation of observed gains into performance metrics. The results can thus be linked to key objectives such as productivity gains, reduction of error-related costs, improvement in service quality, or optimization of time spent on low-value-added tasks.

This final phase thus complements the analyses conducted in the previous stages. While the Discovery phase helps to understand user needs and the controlled experimentation phase analyzes human–AI interaction mechanisms, real-world evaluation measures the solution’s ability to deliver lasting value to users and the organization. It provides the necessary insights to guide decisions regarding rollout, improvement, or investment.

Conclusion

The rapid adoption of generative AI by the general public has profoundly transformed the digital landscape. However, this momentum does not guarantee its successful implementation at the organizational level. Many companies are now finding that technically advanced AI does not necessarily deliver the expected benefits once it is put to the test in real-world work environments.

In this context, evaluating generative AI cannot be limited to measuring its technical performance. Understanding its true value requires analyzing its uses, interactions, and user perceptions, as well as the transformations it brings about in business activities, organizations, and professions.

The method presented in this article is part of this approach. Inspired by the principles of the Living Lab, it proposes a multidimensional approach structured around three complementary levels of analysis: understanding needs and actual work, studying human–AI interactions in a controlled environment, and evaluating impacts under real-world conditions. This approach aims to move beyond a technology-centric view and place the user and their activities at the heart of the evaluation process.

More broadly, we hypothesize that the main challenge in the coming years will not lie in improving model performance, but in our ability to understand the conditions that foster their sustainable adoption and value creation within organizations.

This method is currently being implemented as part of several generative AI evaluation projects with our clients. An upcoming article will present the results of these experiments and analyze the lessons learned from the field, in order to concretely illustrate the contributions, limitations, and prospects for the evolution of this approach.

Bibliography

  • [1] : Simon Borel, Yann Ferguson, Jean Condé, Etude des impacts de l’IA sur le travail. Synthèse générale du rapport d’enquête du LaboIA Explorer. 2024
  • [2] : Yann Ferguson. 1. Ce que l’intelligence artificielle fait de l’homme au travail. Visite sociologique d’une entreprise. Chapitre d’ouvrage : Les mutations au travail p 23 – 42. 2019.
  • [3] : Robin Héron, Myriam Fréjus. Au-delà des discours, l’IA générative à l’épreuve des usages réels en entreprise 1 EDF Recherche & Développement, SEQUOIA. Conference: APIA – Conférence Nationale sur les Applications Pratiques de l’Intelligence Artificielle – Dijon. 2025
  • [4] : Valérie Michel-Pellegrino – Expérimenter l’innovation en situation réelle : clé essentielle de la co-conception centrée utilisateur. Article interne Drit Berger-Levrault, 2025.
  • [5] : Alexandre Agossah, Frédérique Krupa, Matthieu Perreira Da Silva, Guillaume Deconde, Patrick Le Callet. Déploiement de l’IA en situation de travail : une trop faible considération de l’expérience des employé×es ?. Sciences du Design, 2023, n° 16 (2), pp.68-85. ⟨10.3917/sdd.016.0068⟩. ⟨hal-04139664⟩
  • [6] : Alexandre Agossah. Acceptabilité de l’Intelligence Artificielle en contexte professionnel : facteurs d’influence et méthodologies d’évaluation. Intelligence artificielle [cs.AI]. Nantes Université, 2024. Français. ⟨NNT: 2024NANU4022⟩. ⟨tel-04826567 [7] : Living Labs

More ...

Scroll to Top