//pragmatic leaders
← cases

facebook — data, privacy, and the tension at the core of a social network

Facebook's business depends on knowing more about you than you'd consciously share. The product that generates this intelligence is the social graph itself. A case in data ethics, advertising architecture, and the PM's role in navigating them.

The product and the data flywheel

Facebook launched in 2004 as a college social network. By 2019, it had 2.4 billion monthly active users globally, with market capitalisation approaching $600 billion. The company's revenue in 2019 was $70.7 billion, almost entirely from advertising.

The product that generated this value is not the News Feed, the Like button, or the Groups feature. Those are surfaces. The product is the social graph — the mapped network of connections, shared interests, declared attributes, and behavioral signals that Facebook accumulates for every user. The social graph is what Facebook sells to advertisers. The social media platform is the mechanism through which users generate and maintain the social graph.

Understanding Facebook as a product requires holding this inversion clearly: the users are the input, not the customer. The advertisers are the customer. The "product" from a revenue perspective is the predictive intelligence about user behaviour that the social graph enables. Every design decision Facebook makes is ultimately an optimisation of that intelligence — more engagement, more data, more precise targeting.

The Decision: advertising or subscription?

In Facebook's early years, the business model question was genuinely open. A social network with dense engagement could have monetised through subscriptions, through commerce facilitation, or through advertising. Google had demonstrated that targeted advertising against expressed intent was an extraordinarily valuable model. Facebook's bet was that targeted advertising against inferred identity and social context could be equally or more valuable.

The distinction matters. Google's advertising model targets intent: the user types "best hiking boots" and sees ads for hiking boots. The intent is explicit. Facebook's advertising model targets identity and social context: the user hasn't searched for anything, but Facebook knows they're a 34-year-old parent in a specific income bracket who has been engaging with outdoor content and whose friends have recently purchased hiking equipment. The targeting is inferential, not explicit.

For advertisers, inferential targeting at Facebook's scale is often more commercially valuable than intent-based targeting. You can reach a user before they've decided to look for your product — reach them at the moment of receptivity that precedes the search. A well-targeted Facebook ad appears to a user who is already predisposed to the offer, in a context (social browsing) where they're in a consumption mindset rather than an active task-completion mindset. Conversion rates for properly targeted Facebook campaigns consistently exceeded those of equivalent display advertising.

The product architecture that enables this: every connection on Facebook — friend, page follow, group membership, event attendance — is a data point about identity and interest. Every Like, every comment, every post share is a behavioral signal. Every pause on a video, every ad impression not clicked, every link opened — these implicit signals are often more commercially valuable than the explicit ones, because they reflect actual behavioral patterns rather than curated self-presentation.

What Worked / What Failed

The advertising model worked at a scale no one had expected. The precise targeting capability Facebook offered was sufficiently differentiated from prior advertising formats that brands shifted significant budget from television, print, and display advertising to Facebook. The combination of demographic precision, behavioral data, and lookalike audience modeling — finding users who behave like your best existing customers — created an advertising product that performed measurably better than alternatives on a cost-per-acquisition basis. This is why Facebook's revenue grew from $153M in 2007 to $70.7B in 2019. The product worked for the customer who paid for it.

What failed, and what became the dominant story of Facebook's second decade, was the gap between what users understood about the data collection and what Facebook was actually doing with it.

The Cambridge Analytica disclosure in 2018 revealed that data from 87 million Facebook profiles had been accessed without explicit user consent and used for political advertising targeting. The mechanics of what happened matter: a third-party developer had built an app that, under Facebook's platform policies at the time, could collect data not just from users who installed the app but from their entire friend network — meaning a user who had never heard of Cambridge Analytica could have their profile data harvested because a friend installed a quiz app. This was not a data breach. It was the platform API working as designed. The problem was that the designed behaviour was inconsistent with what any reasonable user had consented to when they created a Facebook account.

The response to Cambridge Analytica was a tightening of the developer API and increased regulatory scrutiny across jurisdictions. The EU's GDPR, which came into force in May 2018, imposed consent requirements that fundamentally changed how Facebook could collect and use data for European users. The FTC fine of $5 billion in 2019 — the largest ever levied against a technology company — was the US regulatory response to the same pattern of behaviour.

The failure here is not primarily legal or regulatory. It is a product ethics failure: the design of the data collection and API sharing systems was never evaluated against the question "would users, if they fully understood this, consider it acceptable?" Facebook's historical answer to that question was architectural — build consent mechanisms that satisfy the legal standard, not the standard of what users would actually understand. That approach worked until it generated sufficient public awareness to become a political and regulatory issue.

What a PM should take from this

The Facebook case is not a lesson in how to build a successful data business — that lesson is available, but it is not the useful one. The useful lesson is about the relationship between product design and consent, and what happens when that relationship is managed toward compliance rather than toward user understanding.

Every product that collects user data makes implicit promises about what that data is used for. The implicit promise is set by the context in which the data is collected, not by a terms-of-service document that 0.1% of users read. When you post on Facebook, the implicit context is "I'm sharing this with my friends and the public." The explicit mechanism — the API that extended that data to third-party developers who extended it to their users' networks — violated the implicit context while satisfying the legal terms. The gap between those two is the product failure.

The PM question is not "does this comply with our terms of service?" It is "if a user understood exactly how this data is being used, would they consider it a fair exchange for the service they're receiving?" Designing to that standard requires a different kind of analysis than compliance review. It requires the product team to model the user's mental model of the product — what they believe they've consented to — and to evaluate each data collection and sharing decision against that model.

The second lesson is about how engagement optimisation and user trust optimisation can pull in opposite directions. Facebook's most commercially valuable actions were ones that maximised engagement: more time on the platform, more content interactions, more data generated. Some of the mechanisms that maximised engagement — algorithmic amplification of emotionally provocative content, infinite scroll, notification frequency — also generated user harm in ways that took years to become publicly documented. The PM who is only tracking engagement metrics will not see these costs until they become large enough to generate regulatory action or user defection. The metrics you track determine what you optimise for, and what you optimise for determines what product you build.

// scene:

The internal debate at Facebook about the News Feed algorithm's amplification of divisive content is documented in internal research that became public through the Wall Street Journal's "Facebook Files" investigation in 2021. The research showed that Facebook's own data scientists had identified the link between algorithmic amplification of angry, divisive content and real-world harm. The decision made above them — to maintain the engagement-optimised algorithm — is the decision that product ethics frameworks exist to prevent. The case for those frameworks is that the PM team with clear authority and responsibility for user welfare would have made a different call than the business decision-maker optimising quarterly engagement numbers.