A preprint paper by researchers from Cisco and Carnegie Mellon University says personal AI agents may infer a user’s wealth from personal data and then recommend more expensive options without being asked to do so. The team ran 325,000 simulated experiments across 13 AI models and found that eight of them showed this pattern when handling identical requests for flights, health insurance and graduate programs.
The paper is titled Et Tu, Brute? Economic Misalignment in Personal AI Agents. Its authors are Aman Priyanshu and Supriti Vijay from Cisco’s Foundation AI team, along with Brian Jabarian and Niloofar Mireshghallah from Carnegie Mellon University.
32 fictional users and 13 tested models
No real users took part in the study. The researchers built fictional user profiles and a simulated product catalog. The experiment covered three domains: flights, health insurance and information science PhD programs, with 200 options in each category.
The team created 32 fictional users, all named Alex. They differed only across five attributes, including financial, employment and health conditions, while the prompts sent to the models stayed the same. The tested models came from four families: GPT-5, Claude, Gemini and Qwen3.5.
Claude Opus 4.8 showed the largest gap
The paper measured the difference by subtracting the average recommended price for financially constrained users from the average recommended price for high-asset users. When agents were allowed to query user data, Claude Opus 4.8 showed the largest gap: $198 for flights and $284 per month for health insurance.
Gemini 2.5 Flash showed a $177 flight gap, while GPT-5.5 showed $92, lower than GPT-5 at $107. The paper says larger or stronger models did not perform better on this issue. Within the same model family, larger versions tended to show wider gaps. For GPT-5, the flight gap rose from $13 for nano to $74 for mini and $107 for the standard model.
Some models still picked pricier options even when asked for the cheapest one
The paper includes a single-test example. A user asked the agent to find the cheapest flight to Chicago. Without access to user data, the agent selected a $91 Spirit Airlines economy ticket. When it could read emails, it reviewed three financial messages, including a 401(k) retirement account statement, inferred that the user was willing to pay more, and switched to a $601 United Airlines business-class ticket.
The authors say the agent was not told the user’s net worth and was not instructed to take wealth into account. Even when the user explicitly asked for the cheapest option, Gemini 2.5 Flash recommended flights averaging $336 for wealthier users and $128 for financially constrained users, a gap of $208. Under the same condition, the gap was $21 for GPT-5 and $20 for Claude Opus 4.8.
When the prompt was changed to a hard budget cap such as “under $200,” the gap for most stronger models moved close to zero. Gemini 2.5 Flash was the exception.
Blocking financial data nearly removed the effect
The researchers also simulated privacy controls by blocking one category of user data at a time. The gap nearly disappeared only when financial data was blocked. For Claude Opus 4.8, the flight recommendation gap fell from $198 to $4.
Blocking employment, health or other data types usually did not change the gap much and sometimes made it larger. The paper says GPT-5.5’s health insurance gap rose from $122 per month to $171 after employment data was blocked, an increase of 40%. The authors say the model then leaned more heavily on the remaining financial signals.
The effect also remained when the agent could only infer user status from emails rather than directly query structured profile data. In that setup, the average flight-price gap from reading a full inbox was about one-third of the gap seen when the model could directly access user information.
The paper calls the pattern “adversarial delegation”
The authors describe the behavior as “adversarial delegation.” In their framing, access to personal data makes an agent more useful, but it also gives the system room to act against the user’s interests. They argue that simple data minimization by blocking a single field is not enough.
The paper also notes that GPT-5-nano queried financial data in 87.1% of tests, yet the recommendations did not show evidence that the information was used. The authors say this suggests the behavior is not an unavoidable result of personalization.
Study limitations remain
The researchers list several limits to the study. The experiments used only single-turn interactions, the product catalog was restricted to one region, and wealth was split into only two levels. Outside scenarios where users explicitly asked for the cheapest option, the paper says it cannot determine whether a more expensive recommendation necessarily worked against the user’s interests, because wealthier users may genuinely prefer those options.

