All Reports

Harvard Researchers Develop Priva-See System to Demonstrate AI-Mediated Privacy Leakage from Smartphone Data

Bot Mutiny |

A study of 465 participants reveals that LLM-based systems can accurately infer sensitive personal information, including health conditions and political beliefs, by analyzing photos, calendars, and contacts.

Researchers from Harvard University and Carnegie Mellon University have documented how large language models (LLMs) can extract sensitive personal information from smartphone data that users previously considered private or unstructured. The findings were detailed in a preprint paper titled "Privacy Leakage Through AI-mediated Analysis of Smartphone Data," authored by Sarah Radway, James Mickens, and others, as first reported by arXiv.org. The research team built a system called Priva-See to mirror the capabilities of modern adtech companies. The system uses LLMs to parse multimedia files and unstructured text, such as photos, videos, inboxes, and calendars, which were historically difficult for automated advertising systems to analyze. Unlike traditional tracking that relies on structured data like IP addresses and GPS coordinates, this new method allows for the extraction of sensitive insights from the specific content of a user's life. ## The Priva-See Study

The researchers conducted an IRB-approved study involving 465 participants who deployed the Priva-See app on their mobile devices. The app requested standard permissions including access to photos, calendars, reminders, location, and contacts. The system then sent this local data to an LLM for analysis. The system was designed to speculate about sensitive user characteristics and provide explanations for its inferences. According to the paper, Priva-See made invasive and accurate inferences about the personal lives of users and their social circles despite having access to only a subset of the data available on a typical phone. These inferences included speculations regarding a user's health conditions, political beliefs, and specific physical locations. ## Shifting Data Collection Tactics

The paper notes that the online advertising industry is currently facing a "reduced availability of data signals" due to new restrictions from smartphone vendors on unique per-user identifiers. This shift was highlighted in Meta's 2025 Form 10-K submission to the U.S. Securities and Exchange Commission, which identified the loss of these signals as a significant risk to its adtech business. As traditional tracking identifiers disappear, companies are incentivized to mine inferences directly from the data users disclose themselves. The researchers argue that users often don't understand that granting an app access to a photo doesn't just share the bytes of that image; it gives the app access to every inference an AI can draw from the content of that image. ## Industry Breaches and Data Brokers

The study contextualizes these risks by citing documented failures in the current data ecosystem. In 2024, the FTC sanctioned Gravy Analytics for selling location data from 1 billion devices without consent. This data tracked visits to medical offices and places of worship. In 2025, Gravy Analytics suffered a data breach where hackers stole terabytes of location data harvested from popular apps including Tinder, Candy Crush, and pregnancy trackers. Similarly, the paper cites Datamaster for the unauthorized collection of names, addresses, and contact information of millions of individuals with specific medical conditions. The researchers state that the integration of LLMs into this ecosystem makes these existing risks more acute. ## Proposed Changes to OS Permissions

Based on the performance of Priva-See, the authors suggest that smartphone operating systems must change how they gather user consent. Current permission prompts don't inform users about downstream data usage capabilities or the implicit sharing of personal information that occurs when an AI analyzes a photo or a calendar entry. The researchers also recommend that LLM providers update their "acceptable use" policies to prevent the deployment of Priva-See-style inference systems. They warn that their study likely understates the actual risk, as real-life data brokers have access to significantly more computing power and VRAM than the researchers used for their prototype.

References

(2026). Privacy Leakage Through AI-mediated Analysis of Smartphone Data. arxiv.org. https://doi.org/10.1145/2976749.2978313