Beyond the Told

by Dr. David M Robertson

Your Reviews Can Be A Security Risk

review security risk

Most people assume a public review is low-stakes. You rate a restaurant. You write a paragraph about a hotel. You leave a snide comment under a product because it let you down. It feels anonymous enough. You’re not posting your address. You’re not listing your contacts. You’re just saying what you thought. But that assumption is getting expensive.

A 2026 study in Information Systems Research suggests that there are a few things you need to keep in mind, beyond the usual “be careful what you post.” The claim is that public review behavior can actually reconstruct a usable social graph, and that graph raises the expected return of spear-phishing.

In plain English: Yes, you should be careful about posting locations, family members, business info, or any other type of sensitive information. However, where you review and how long you write can help an outsider infer who you’re connected to and reveal details about your personality. And once someone knows how you write, and if they can guess your trusted circle, impersonation gets cheaper and more precise.

So, the danger isn’t the obvious data. It’s the leftover patterns and behaviors that allow people to recreate your persona. If they do that, you would be amazed at what people can get out of them.

What the researchers actually did

The team treated review activity as an action matrix: which businesses a user reviewed, and how long those reviews were. Their algorithm, Homophilous Network Learning (HNL), works backward from those patterns and asks a simple question: what hidden network of ties would best explain the observed behavior?

Ground truth was Yelp’s published friends list. The model never received those friend lists as inputs. Samples included New Orleans (2,065 reviewers, 6,094 businesses) and Pittsburgh (2,234 reviewers, nearly 16,000 businesses). At a 10% false-positive rate, the model recovered about half of actual Yelp ties in both cities. At a 20% false-positive rate, it recovered more than 60%. Older correlation methods and competing models underperformed.

Review length alone was informative enough that the authors treat it as a leakage channel, not throwaway metadata. The key point is that platforms and users tend to treat length as style. Well, the study treats length as signal.

Why a reconstructed graph matters for phishing

For clarity, spear-phishing is a targeted type of phishing where hackers send highly convincing emails to specific people or small groups inside an organization to steal sensitive information, get login credentials, or trick someone into sending money or opening malware. Spear-phishing works when the message feels like it came from someone you already trust. And that’s the danger. The attacker doesn’t need your password first. They need a believable relationship.

Well, the researchers modeled an attacker who used inferred ties to impersonate a trusted contact. Payoff estimates use FBI Internet Crime Complaint Center loss figures. In the Pennsylvania sample, modeled return on attack rose from 109% at 500 impersonation attempts to 1,098% at 10,000.

The mechanism is precision. Better guesses about who knows whom raise hit rate. Hit rate scales payoff faster than campaign size alone. So, you’re not just sending more spam. You’re aiming better. Hence the term “Spear-phishing.”

The real privacy problem

People usually think privacy fails when a platform publishes a friend list or a contact book. That’s not always true, and this work proves the point. The privacy problem is the behavioral residue platforms treat as harmless when they ship public datasets or APIs.

Think of ratings, comments, and review-style activity. The same class of leakage can apply wherever platforms publish that kind of trail. Think about it: Groupon, Amazon, and YouTube. Do you let your personality show? Do you describe use-cases? I know I have at times.

The point is that even if you scrub “sensitive fields,” leaving the action matrix intact may still release the raw material of a social graph. That’s sometimes more valuable than anything else, because those are the cues that people often listen to when establishing truth and trust. It’s messy!

What they propose to blunt it

The paper proposes two differential-privacy mechanisms. These include direct and indirect. Both are interesting.

  • Direct: add Gaussian noise to the review-length matrix.
  • Indirect: perturb the inferred network, then resynthesize the released data with Laplace noise.

That may not mean anything to you, but in their tests, about three-quarters of review lengths stayed unchanged, and 95% of entries shifted by roughly 15 words or less. That was enough to push smaller campaigns into negative return and to cut profits on larger ones, while keeping most of the data’s analytic use. The indirect route generally confused the attacker more.

Translation for use: you don’t have to destroy the dataset to degrade the phishing input. You have to stop treating review length and related behavioral fields as harmless. They’re not. In other words, if you’re going to leave a review, don’t provide your life story to do it. Keep it simple, stale, and direct.

What you can actually do now

Here’s the hard part for individuals: your options are limited. You can’t usefully noise your own historical review length after the fact, and you can’t privately rewrite the public matrix a platform already released. But what you can do is be more intentional about what you publish going forward, and you can treat unexpected “trusted contact” messages with more suspicion when the relationship cue feels convenient.

Ultimately, I guess the lesson is that ordinary public acts leak social structure, but social structure is the raw material of impersonation. So, if your security model only watches for addresses, passwords, and friend lists, you’re only doing part of the work. And with AI, impersonation is getting extremely sophisticated. Just something to think about.


If you got something from this article, you might also like AI Has Rewired Fraud. Here’s What Works.

Source for more research: Leng, Y., Chen, X., Dong, X., Wu, L., & Shi, Z. (2026). When behavioral data betray users: A diagnostic and protective framework against social interaction leakages. Information Systems Research. https://doi.org/10.1287/isre.2024.1469 Working paper: SSRN 3875878 Popular write-up: StudyFinds, 8 Sep 2026