Skip to content
  1. Home
  2. Blog
  3. AI User Engagement: The Problem Personalization Can’t Solve
Behavioural Science

AI User Engagement: The Problem Personalization Can’t Solve

Joris Beerda16 June 2026
AI User Engagement: The Problem Personalization Can’t Solve


Microsoft Employees in their support division had access to one of the most capable AI systems ever deployed inside an enterprise, and most of them barely touched it. The ones who did CoPilot it treated it like a search box: type a question, copy the answer, close the window. There was no journey to return to. Nothing connected one session to the next, and nothing made the tenth interaction feel any different from the first.

AI is often heralded as the end of the engagement problem, because it supposedly solves the personalization of every user experience. The reality of AI user engagement is close to the opposite. Capability is astonishing, but the experience of using these tools remains technical, linear, and forgettable. People try them, fail to form a habit, and quietly stop.

A working definition before we go further. AI user engagement is sustained voluntary adoption with deepening capability: people choose the tool when nothing forces them, their use grows more skilled over time, and they would object if you took it away. Session counts and provisioned seats measure exposure. Engagement is what happens even after the initial novelty wears off.

The data backs up what we saw in the field. MIT’s NANDA initiative reported in 2025 that about 95% of enterprise generative AI pilots and projects deliver no measurable P&L impact, with only around 5% reaching production and driving real revenue. Gartner predicted in June 2025 that half of the organizations planning to cut customer service staff in favor of AI will abandon those plans, and in February 2026 went further: by 2027, half of the companies that did cut staff will be rehiring for similar roles. Stack those findings and the story is consistent. AI is everywhere, yet little of it is changing behavior the way its buyers assumed.

After two decades of designing engagement for companies like Microsoft, Porsche, and Booking.com, we believe the diagnosis is simpler than the industry wants it to be. AI products have a motivation problem, and no amount of personalization fixes it.

The model knows what you clicked, bought, or asked. It does not really know why, and without the why, so it is . Until that changes, every additional dollar of model spend, license expansion, and feature launch keeps landing on the same flat usage curve.

What Is AI User Engagement?

AI user engagement is the sustained, voluntary use of an AI product in which people grow more capable over time, choose the tool when they have alternatives, and would object if it were taken away. It is distinct from adoption or exposure metrics, which count access rather than genuine pull. In practice it shows up as three things:

  • Voluntary return: people come back without a mandate or a notification prompting them.
  • Deepening use: the same person’s usage broadens and grows more skilled week over week, rather than plateauing after the first session.
  • Felt ownership: users accumulate work, history, or standing inside the product that they would not want to lose.

Everything below is about why those three are so rare in AI products, and how to design for them. The distinction matters because most organizations measure exposure (seats provisioned, prompts sent) and mistake it for engagement. Exposure tells you who has access. Engagement tells you who would care if you took it away. And in the end that is what matters. Technological innnovation only works if humans seek to interact with it continuously and over time.

The Pattern We Saw Inside Copilot

When organizations roll out an AI assistant, the implicit assumption is that capability creates engagement. The tool can draft your emails, summarize your meetings, and analyze your spreadsheets, so naturally people will use it every day. What we observed in the Copilot work is that this assumption fails for most employees.

The search-box behavior is the tell. Users asked Copilot a question the way they would use Google, took the answer, and left. Nothing in the interaction acknowledged who they were, what they had done last week, what other people were searching for, or achieving, or what they were getting better at. Every session started from zero.

In Octalysis terms, the experience did not evolve with the seniorty of the user in the experience. No Onboarding loop that delivers a first major win state. No Scaffolding that carries a user toward mastery. No Endgame that gives a veteran something to become. Each session restarted at zero, and zero is where most users quietly stayed.

Compare that to any product people return to daily without being reminded, like Duolingo. There is always a reason to come back that exists before the need arises: a streak to protect, a progression to continue, a community to check in on, a craft to refine.

Copilot, like most enterprise AI deployments, offered none of these. The only thing that brought a user back was their own memory that the tool existed, which is not an engagement strategy so much as a hope dressed up as one.

McKinsey’s Superagency research hints at how much room this leaves on the table. Roughly three-quarters of employees report using AI in some capacity at work, yet C-suite leaders estimate that only 4% of employees use generative AI for a third or more of their daily tasks, while employee self-reports put the real figure around three times higher. Employees are experimenting far more than their leaders realize; what’s missing is the structure that turns scattered experimentation into compounding capability. That gap between “tried it” and “works with it” is the gap engagement design exists to close.

Why Workplace AI Rollouts Stall

A mandate buys you compliance; engagement is what buys you capability. Most AI rollouts confuse the two, and the confusion is expensive.

The arc repeats across industries. Licenses are bought, training sessions are scheduled, a usage dashboard goes up, and within a quarter the dashboard becomes the uncomfortable slide in the steering committee deck. Leadership responds the way leadership responds to any flat metric: usage becomes a KPI.

And usage duly rises, because employees perform any behavior that is audited. Opening the tool to satisfy a target is Loss & Avoidance (avoiding the ire of management) doing the work, and Loss & Avoidance reliably produces the minimum viable behavior and makes people feel out of control. People paste one paragraph into the assistant, collect their checkmark, and go back to working the way they always have. The metric improves while the capability it was supposed to measure never develops.

What separates the rollouts that take from the ones that stall is rarely the tool. McKinsey’s Superagency data points at the social layer: employees adopt AI productively when their direct manager actively champions it as a way for the team to grow. Frame the same deployment as monitoring or as a prelude to headcount cuts, and you have flipped the motivational valence of every interaction from Empowerment of Creativity & Feedback to Scarcity & Impatience and fear. Identical software, opposite Core Drives, opposite outcomes.

Gallup measured global employee engagement at 20% in 2025, its lowest level since 2020. An AI assistant dropped into a disengaged workforce inherits that disengagement. And the adoption-to-value gap is now visible at every altitude: McKinsey’s State of AI survey finds the large majority of organizations using AI in at least one function, while only a small minority report meaningful enterprise-level earnings impact from it. Nobody explores a new tool with curiosity when they feel no curiosity about the work itself, which is why we keep telling clients that an AI rollout is an engagement program wearing a software license, whether anyone designed it that way or not.

Voluntary adoption can be designed. In our own client work on the gamified sales program for P&G’s distributor network, the rollout was non-mandatory and reached 99.5% adoption, because participating was made more interesting than not participating. It sets the standard we hold AI deployments to. If a tool this capable needs a mandate to get used, the experience has failed its Discovery phase, underwhelms in it Onboarding, and stalls during Scaffolding. No amount of training calendar can compensate.

Gamification Already Taught This Lesson

The AI industry is rerunning an experiment whose results are already in. In 2012, at the peak of the first gamification hype cycle, Gartner predicted that by 2014, 80% of then-current gamified applications would fail to meet their business objectives, primarily because of poor design. The prediction held up. Companies had bolted points, badges, and leaderboards onto experiences nobody had designed motivationally, and users saw through it within weeks.

Swap the nouns and the sentence describes the present. Companies are bolting chatbots and copilots onto experiences nobody has designed motivationally, and MIT’s 95% pilot failure rate is the same finding wearing newer clothes. What changed was the technology, not the mistake.

We watched the first cycle from the inside, and the pattern of who survived it is instructive. The gamification work that outlasted the trough treated mechanics as surface expressions of underlying motivation, while everything that treated the mechanics as the product died on schedule. A point is worthless until it marks progress someone cares about, and an AI response is equally worthless as a retention device until it sits inside an experience someone wants to continue.

The same selection event is coming for AI products. When every competitor runs a frontier model, capability stops differentiating, and the remaining question is whose experience people choose when nobody is forcing them.

The 29-Minute Paradox

Here is the comparison that should unsettle every product team bolting AI onto their experience. The most engaging AI products in the world are not the most capable ones.

When Character.AI launched its mobile app, users were averaging 29 minutes per visit, a figure the company claimed eclipsed ChatGPT’s by 300%. Its investor Andreessen Horowitz reported that active users spend around two hours a day on the platform. The underlying model was never close to frontier quality, and its users did not care. They were not there for answers. They were in an ongoing relationship with characters they had often built themselves, inside conversations that accumulated history and meaning.

Meanwhile one of the most capable general assistants on the market gets used, by OpenAI’s own account, mostly as an answer machine. The company’s first large-scale usage study, published with NBER in September 2025, found that practical guidance, information seeking, and writing dominate ChatGPT conversations, and that the clear majority of usage is personal rather than professional, a share that had been climbing year over year. Ask, receive, leave. People clearly like the tool, casually and personally, without it becoming part of any structure in their lives. The study describes a utility, and utilities are judged on speed, which means the moment a faster answer source appears, the user follows it.

Read those two data points through a behavioral lens and one conclusion follows: engagement lives in the interaction design, not in the model. People come back to where they feel ownership, progress, and a sense of belonging, none of which a sharper answer supplies. Character.AI’s users return because the experience activates Ownership & Possession (my characters, my conversations) and Social Influence & Relatedness (a relationship, however synthetic).

Duolingo proves the same point from the opposite direction, with AI serving design instead of replacing it. Birdbrain calibrates the billion-plus exercises learners complete daily, yet what carries its more than 55 million daily active users is the streak, the leagues, the progression. In Duolingo’s own A/B testing, offering a streak wager lifted day-7 retention by 14%. The AI decides when and what. The behavioral design decides why.

One caution before anyone treats companion AI as the blueprint. Two hours a day of synthetic companionship is a different kind of product with different aims, and a serious productivity or enterprise tool will not generate that pattern of use, nor would it want to. The transferable lesson from Character.AI is structural: a designed journey makes people return because they grow, create, and connect, and that is the part worth copying.

What the Sticky AI Experiences Share

Look across the small set of AI experiences with real retention and three patterns repeat. None of them depends on having the best model.

First, the engagement loop existed before AI did. Duolingo did not become sticky when Birdbrain arrived. The streaks, leagues, and progression systems were already carrying daily behavior, and the AI was slotted inside that loop to make each step better calibrated. The loop creates the reason to return; the model improves what happens once you are there. Products that reverse the order, shipping a model and hoping a loop emerges, are the ones writing the retention curves that MIT’s failure numbers describe.

Second, users accumulate something they own. Character.AI users build characters and conversation histories they would feel a real loss abandoning. The same demand surfaced inside ChatGPT itself when OpenAI let users build custom GPTs: over 3 million were created within two months, before the store even launched. That is Ownership & Possession and Empowerment of Creativity & Feedback announcing themselves at volume. Users were begging for a way to invest in the product rather than merely query it.

Third, the AI sits inside the moment of work rather than waiting in a separate destination. Coding assistants succeed disproportionately because they live in the editor, where the user already is, attached to a task the user already cares about, inside a craft the user is already motivated to master. A separate tab with a blank box asks the user to bring their own motivation. An embedded assistant borrows motivation the workflow already generates.

Loop first, ownership second, embedding third. Hold any stalled AI deployment against those three and the failure usually localizes fast.

AI Products Through the Four Experience Phases

Every Octalysis engagement design runs through four experience phases, because motivation is not one thing applied evenly. It changes shape as a person moves from stranger to veteran. The phases are Discovery, Onboarding, Scaffolding, and Endgame, and each carries a distinct motivational job. Map almost any AI product against them and the same diagnosis appears: the technology is exponential, the experience design is a flat line.

Discovery: why would anyone start?

Discovery happens before a person is ever inside the experience. It is the phase that answers a single question: why would someone want to begin in the first place? For most AI products, the honest answer is “because everyone says I should,” which is borrowed motivation, not designed motivation.

The pitch is capability (“it can do anything”) rather than a reason that connects to what a specific person cares about. Capability impresses without compelling. A tool sold as infinitely capable gives a prospective user nothing concrete to want, which is why so many AI rollouts begin with a spike of curiosity and no second act.

Onboarding: the first real win

Onboarding begins after sign-up, when the user runs their first activity loop. Its job is to make the person feel competent and certain they have come to the right place, and to deliver a first genuine win that creates the curiosity to return.

This is exactly where most AI products collapse. They greet a new user with an empty prompt field and infinite possibility, which behaviorally is indistinguishable from a blank page. Infinite possibility with zero direction produces paralysis, and paralysis produces the most common AI usage pattern we see inside client organizations: one ambitious first session, two mediocre follow-ups, then silence. The user concluded the tool “doesn’t really work for me,” when in truth nobody designed their first win. Prompting is a real skill with a real learning curve, and almost no product treats that first loop as something to be guided toward a success the user can feel.

Scaffolding: the shift to mastery

Scaffolding is the long climb toward mastery, and it carries the most important motivational move in the whole model. Early on, a user is reasonably driven by left-brain, extrinsic Core Drives: rewards, progress markers, the avoidance of falling behind. Those work for a while, then curdle into burnout or boredom if they keep leading. So in Scaffolding the design deliberately shifts weight toward the right-brain, intrinsic drives: Empowerment of Creativity & Feedback, Social Influence & Relatedness, the meaning of the work itself. The extrinsic drives stay in the mix, but they stop leading, because nobody builds a lasting habit out of pure reward-seeking.

Current AI products never make this shift, because they never built the extrinsic scaffolding to begin with. Session fifty looks exactly like session one. Nothing the user did has deepened into anything they are now better at or more invested in.

Endgame: turning veterans into ambassadors

Endgame is for the people who have been inside the experience a long time and have done nearly everything there is to do. The question is what remains for them, and the answer that builds durable advantage is a role: veteran users become internal and external ambassadors, the people who teach newcomers, shape the community, and carry the brand outward on their own initiative.

What does a veteran AI user become today? Nobody. A user’s thousandth prompt earns exactly what their first one did, with no status to hold, no body of work to stand behind, and no community to lead.

The GPT Store episode shows how close the industry came to noticing. Three million user-built GPTs in two months was the raw demand for an Endgame: people wanted to create, publish, and be recognized. The supply side never followed. Builders got no progression, no meaningful status, and a revenue program that stayed an afterthought, so the ambassador energy dissipated. The first AI platform that treats its power users as a community to develop rather than a usage statistic will discover what game designers have known for decades about where durable engagement comes from. We have seen this play out across 175+ client engagements: the endgame is where most products leave the most value on the table.

The Chatbot Fallacy

Almost every client conversation about AI eventually arrives at the same request: put a chatbot into the user experience. The assumption underneath is that conversational access to an LLM is itself engaging. It is not, and the most expensive public experiment in this space proved it.

Klarna automated the majority of its customer service chats with an AI assistant that, by the company’s own account, handled 2.3 million conversations in its first month, work it equated to roughly 700 full-time agents. The efficiency numbers were spectacular and widely promoted. Then CEO Sebastian Siemiatkowski publicly conceded the pivot had produced lower quality service, and the company began recruiting humans back into the loop. The bot could resolve tickets. What it could not do was read what a frustrated customer with a disputed charge actually needed, which was sometimes a refund, sometimes reassurance, and sometimes just to be taken seriously by someone with authority.

That failure generalizes. An LLM dropped into a user experience knows the content of the current conversation and, at best, a transaction history. It does not know whether this user is motivated by achievement or by belonging, whether scarcity excites or repels them, whether they want efficiency or want to explore. So it speaks in the same competent, compliant, slightly servile register to everyone, and users feel the genericness immediately.

Twilio Segment’s State of Personalization research has tracked this perception gap for years: in its reporting, around seven in ten consumers expect personalized experiences and roughly three-quarters get frustrated when they don’t receive them, while companies rate their own personalization far more highly than their customers do. AI chatbots are now recreating that same gap at higher resolution: ever more detailed content personalization, with no real model of what the user needs from the interaction.

Service design through a motivational lens looks different. A customer disputing a charge is deep in Loss & Avoidance; what de-escalates them is restored control and a credible path to resolution, not a cheerful apology template. The confused first-time user needs competence built up in small steps, not a link to documentation. And the loyal customer reporting a minor bug is often expressing investment in the product, so the right response strengthens the relationship rather than merely closing the ticket.

An AI that could read which of these states it is facing, and shift register accordingly, would deserve the word personalization. Klarna’s eventual hybrid model, with AI handling routine queries and humans taking over where empathy and discretion decide the outcome, is this insight rediscovered manually and priced in headcount.

The industry’s standard answer to the knowing-the-user problem is to gather more declared data: onboarding quizzes, preference surveys, periodic check-ins. Behaviorally, this is a tax. Surveys interrupt the experience, answers decay within weeks, and people are famously unreliable narrators of their own motivations. Asking users to explain what drives them is outsourcing the hardest design problem to the people least equipped to answer it.

Personalizing Content Is One Thing. Personalizing Motivation Is Another.

The deepest confusion in the AI engagement conversation is treating these as the same problem. They are different problems, and the second one is where retention is decided.

Content personalization answers: what should this user see next? Recommendation engines have done this well for a decade, and LLMs do it better. Motivation personalization answers a harder question: what should this experience feel like for this user, and what should it offer them a chance to do?

Two users can perform the identical action, abandoning the same flow at the same step, for opposite motivational reasons. One left because the experience offered no challenge. The other left because it offered no connection. Serve both of them the same smart re-engagement message and you will win one back while pushing the other further away.

Picture it concretely in a fitness app. Two users skip three workouts in a row, so the behavioral data is identical. User one is bored: the program stopped challenging her weeks ago, and what would bring her back is a harder goal and a way to measure herself against it.

User two is isolated. She joined because her friends were on the app, the group went quiet, and what would bring her back is a partner challenge, never a more difficult workout plan. Content-level AI sends both women the same “We miss you! Here’s a personalized plan” message, and motivation-level AI would not even pick the same channel of intervention, let alone the same words.

Every system we audit, from loyalty programs to enterprise platforms, eventually runs into this wall. Here is the wall stated plainly: AI knows what you did, and it does not know why you did it. It can see that you bought the product, clicked the button, and abandoned the cart, and from that history it builds a guess about who you might be. What it cannot see is the motivation underneath the click.

Were you driven by Social Influence & Relatedness, doing what people you admire are doing? By status and the visible progress of Development & Accomplishment? By Loss & Avoidance, the fear of missing out or falling behind? The same purchase can come from any of those drives, and each one calls for a different next move.

An AI optimizing against behavior alone gets superb at predicting what a user does and stays blind to what would make that user care. This is why personalization metrics keep improving while engagement metrics keep falling, in workplaces and in consumer products alike.

The motivational layer is exactly what the Octalysis Framework was built to describe. Created by Yu-kai Chou, whom industry peers and gamification rankings routinely name among the field’s foremost authorities, the framework has become one of the most widely taught and applied models of human motivation in the world. Its question is the one personalization keeps skipping: which Core Drives is this experience activating, for this user, in this phase of their journey? Answer that, and AI personalization finally has a target worth optimizing toward.

Inferring Motivation: Where AI User Engagement Goes Next

The question we kept running into, with Microsoft and with every client who wants a chatbot, is the one that motivated years of our own research: can a system understand what motivates a user without ever putting them through a survey? Surveys are the industry’s reflexive answer, and they fail twice over. They tax the experience, and people are unreliable narrators of their own motivation. The answer had to come from behavior itself.

It can, because motivation leaks into behavior constantly. A user who reorganizes their workspace, customizes settings, and curates collections is expressing Ownership & Possession without being asked. One who opens every new feature announcement is expressing Unpredictability & Curiosity. One who shares results and joins every group activity is telling you Social Influence & Relatedness carries them.

The signals are already in the interaction stream. What has been missing is a model of human motivation rich enough to read them, which is exactly what two decades of Octalysis work across hundreds of engagements produced.

This is the breakthrough we built, now patent-pending, as the Octalysis AI Engagement Engine. It treats motivation as a probability problem. It ingests the behavioral exhaust a product already generates, things like login cadence, the kinds of prompts a person writes, session length, which features they return to, what they customize, and how they interact with teammates, and maps those signals onto the eight Core Drives. From that stream it assigns and continuously updates probabilities: how likely is it that this person is driven right now by status, by relatedness, by the fear of loss, by curiosity? As evidence accumulates the probabilities sharpen, and the system converges on an accurate read of who a user is and what moves them in each phase of their journey, never once interrupting them to ask.

Then it does the part that matters. It adapts the motivational structure of the experience, the challenges, the social surface, the rewards, the narrative, not merely the content on the screen. The chatbot stops being a generic oracle that talks to everyone the same way and becomes an experience that knows whether you need a harder challenge or a warmer check-in, without ever having asked.

The phase dimension is built into the model, because a single read is not enough. The socially driven user who joined for her friends may, four months in, have developed a real appetite for measurable progress. A static profile would keep serving her partner challenges long after they stopped landing. Motivation shifts as competence grows and novelty fades, and the Engine is built around that evolution rather than a one-time snapshot.

This is the layer the largest AI companies are missing. The technical personalization at Facebook, Google, and every frontier lab is extraordinary at modeling behavior and weak at modeling motivation, because behavioral data tells you what happened and never why. Pairing that technical capability with a rigorous model of human motivation is the next leap, and it is the leap we have spent twenty years preparing to make. The companies that get there first will hold an advantage better models alone cannot erode, because their users will have a reason to stay that has nothing to do with benchmark scores.

Is Engagement Even the Right Goal for AI?

An honest objection deserves an answer before the recommendations. Plenty of AI is infrastructure, and infrastructure should be invisible. A fraud-detection model needs no streak, and an assistant whose whole job is to make a task disappear should not be optimized for session minutes. Optimizing for raw time-in-product measures the wrong thing entirely, rewarding minutes spent rather than capability gained.

So the definition from the top of this article matters: sustained voluntary adoption with deepening capability. For a workplace assistant, that looks like an employee whose tenth week of usage is broader and more skilled than their first. For a consumer product, it looks like a user who returns unprompted because the experience has become partly theirs.

Joris Beerda is CEO and Co-Founder of The Octalysis Group, the behavioral design consultancy behind engagement programs for Microsoft, Porsche, Booking.com, LATAM Airlines, and CAIXA. He is co-inventor of the patent-pending Octalysis AI Engagement Engine. See more of the firm’s work in its client case studies.*

Frequently Asked Questions

What is AI user engagement?

AI user engagement is sustained voluntary adoption of an AI product with deepening capability over time. People choose the tool when nothing forces them, their use grows more skilled across weeks and months, and they would object if it were taken away. It is distinct from exposure metrics like provisioned seats, prompts sent, or weekly active users, which measure access rather than genuine engagement.

Why do most AI pilots and rollouts fail?

The common cause is missing motivation design, not model quality. Most AI products offer capability without a designed experience: no first win in onboarding, no path to mastery, no reason to return before the need arises. MIT’s NANDA initiative found about 95% of enterprise generative AI pilots and projects deliver no measurable P&L impact. A capable tool with no designed journey gets used occasionally and abandoned quietly.

What is the difference between content personalization and motivation personalization?

Content personalization decides what a user should see next, which recommendation engines and LLMs already do well. Motivation personalization decides what the experience should feel like for a specific user and what it should offer them a chance to do. Two users can take the identical action for opposite reasons, one seeking challenge and the other seeking connection, so behavior data alone cannot tell you how to respond.

Can AI understand what motivates a user without surveys?

Yes. Motivation expresses itself in behavior: customizing a workspace signals Ownership & Possession, joining every group activity signals Social Influence & Relatedness, chasing visible progress signals Development & Accomplishment. The Octalysis AI Engagement Engine reads these signals as a probability distribution across the eight Core Drives and sharpens it as evidence accumulates, inferring motivation without asking the user to fill in a form.

How do you measure AI user engagement?

Track behavior that reveals genuine pull rather than mandated activity: the share of users who return without any notification, whether an individual’s usage deepens and broadens week over week, and how long a new user takes to reach their first felt success. These predict whether an AI investment survives far better than seat counts or total prompts.

Joris Beerda

Co-Founder and CEO of The Octalysis Group. As a world-leading expert in Human-Focused Design and Octalysis Gamification, Joris’ global career in creating engagement spans across 20 years, 15 countries and 7 languages. He has designed Human-Focused experiences for dozens of Fortune 500s as well as medium sized companies. Joris is also a well known Keynote Speaker on Gamification in many renowned conferences throughout Europe, Asia, and Australia.

Featured case study

Leave a comment

Your email address will not be published. Required fields are marked *

Already have a project in mind? Let’s work together!

Our team will listen to your engagement challenges and decide if we are the best match.

Get in touch

Case studies

Measured results from our client work

See all 14 case studies

Education And Technology

Gamified Learning Platform

Gamified learning platform with clear learning journeys, deepening learning outcomes and reducing reliance on first-line support functions.

90%+Learners actively progressing

Loyalty

Gamified Loyalty Program

Gamified loyalty experience with personalized missions and rewards, leading to record-breaking engagement and conversions.

+153%Increase in credit card acquisitions

Fast-moving Consumer Goods

Sales Force Engagement

Gamified sales platform transforming reps into virtual traders, boosting engagement and performance across 19 countries. Best Gamification Award.

+28.6%Increase in sales revenue