How Companies Monetize Personal Data and Its Impact on Privacy

•

 18 min video

•

 8 min read

YouTube video ID: 6BOxK_JrghY

Source: YouTube video by How Money Works — Watch original video

PDF

Big companies are increasingly collecting and leveraging personal data for profit, a trend that has become more pervasive. While most people accept that tech giants like Google, Facebook, and Amazon gather data for targeted advertising, the scope of data collection has expanded significantly. Now, cars monitor driving habits, smart fridges track eating patterns, smartwatches monitor health, and smart TVs observe viewing habits. Even DNA sequences are part of an active marketplace. This raises questions about the financial viability of such extensive data harvesting, with estimates suggesting companies need to extract around $2,000 per user annually. This implies that for some individuals, their data might be worth more than they are.

The traditional adage, "if you are not the customer, you are the product," no longer fully captures the current landscape. In many new markets, individuals are not even the product; they are merely the raw materials.

The Evolution of Data Collection

The concept of "big data" is relatively new, but the practice of collecting personal information for commercial purposes has a long history.

Early Data Collection: Equifax's Predecessor

Retail Credit Company, later renamed Equifax, was founded in 1899. For much of its early history, it created "fact cards" about individuals, which were sold to insurance firms and lenders. These cards contained not only financial information but also non-standardized details, including rumors about marital troubles, sex lives, and political leanings, which were considered relevant for credit applications.

From the company's perspective, these records were expensive to produce and required significant human discretion to interpret. Bank managers at the time were highly valued professionals because their experience in assessing borrowers was crucial to a bank's profitability. Due to the cost, data collection was primarily limited to high-income individuals.

The Digital Transformation and Standardization

Around the 1970s, companies like Retail Credit began computerizing their records. This coincided with the passage of the Fair Credit Reporting Act, leading to Retail Credit's rebranding as Equifax. Digital records offered several advantages: - Easy retrieval: No more sifting through paper files. - Cost-effective copying: Digital copies were virtually free. - Standardization: Data could be uniformly organized.

In 1989, Fair Isaac Corporation launched the FICO score, which aggregated various data points into a single number. By 1995, the FICO score became an industry standard, adopted even by major mortgage guarantors like Fannie Mae and Freddie Mac. This standardization significantly reduced costs for the financial industry, allowing for large-scale lending without as much individual discretion from bank managers. While this made loans cheaper and more accessible, it also led to the automation of many bank manager roles.

Why Data Collection Exploded

The shift from narrow financial applications to widespread data tracking is due to several factors:

Decreased Cost of Data

  • Storage: In 1981, a gigabyte of hard drive space cost hundreds of thousands of dollars. Today, it costs less than a dime. This drastic reduction in storage costs made it feasible to collect and store vast amounts of data.
  • Collection: Modern connected cars can generate terabytes of raw data daily. Such volumes would have been impossible to store decades ago, making it financially unviable to track granular details like driving habits for minor premium adjustments.
  • Analysis: Early computers were less powerful, making data analysis slow and difficult. Modern computing power allows for rapid and sophisticated analysis, enabling companies to fine-tune even mundane details like optimal times for targeted ads.

Cheap storage also facilitated data trading, and since copying data was free, companies began hoarding it as a strategic investment.

Beyond Obvious Applications: The Hidden Value of Your Data

While targeted advertising, credit scoring, and insurance risk adjustment are well-known uses of data, companies have found less obvious and more impactful applications.

The Employee Score: The Work Number

Equifax offers a service called "The Work Number," a database similar to a credit score but for employment history. Employers submit information such as: - Start and end dates - Income - Insurance status - Pay frequency - Job title - Unemployment claims after leaving - Most critically, whether a previous employer would rehire the individual.

This database contains over 839 million employee records from more than 5 million employers. A negative mark, like being fired or quitting on bad terms, can follow an individual, similar to a loan default. Unlike credit reports, which have rules for negative marks to expire, protections around employment data are much weaker.

This data is valuable even for exemplary employees who were laid off, as visible unemployment can reduce negotiating power. In the past, individuals could fudge employment details, and former managers might provide positive reviews. However, with databases like The Work Number, information asymmetry heavily favors employers.

The impact is tangible: when states banned employers from asking about salary history, pay for job changers increased by about 5%. With this data, employers no longer need to ask. It also helps them sort through applications to identify candidates with the least negotiating power, often using automated screening software to find those who have been out of work longest or had the lowest previous salaries. Equifax openly sells "pre-hire verification" services using this database. This demonstrates how data can derive value from an individual's ability to earn an income, a much broader pool than discretionary spending.

Other Creepy Uses of Big Data

  • Dynamic Pricing: Algorithms use data on income, occupation, marital status, and past shopping habits to determine the maximum price a customer is willing to pay for a product at any given time.
  • Quantitative Investment: Investment firms hoard vast amounts of "alternative data" (credit card receipts, satellite photos of parking lots) to feed into models for better risk-adjusted returns. This is a $30 billion industry.
  • Tenant Scoring: Landlords use similar data points to score potential tenants.
  • Automotive Data Sales: Car companies sell driving habits to insurance firms. GM was caught passing driver behavior data from connected cars to LexisNexis and Verisk, which then packaged it for insurers. One Chevy Bolt owner saw his insurance jump 21% after his driving data, including 640 logged trips, was shared. Honda reportedly received only 26 cents per car for handing over customer driving data.
  • Political Campaigns: Campaigns compile household-level data for individualized messaging, matching voter households to IP addresses to target every device up to three times a day.
  • Law Enforcement: Taxpayer-funded law enforcement agencies extensively use data collection and analysis.

This data collection goes beyond selling individuals as products to advertisers; it leverages their votes, rent, livelihoods, and even their potential to become crime statistics.

The Data Flywheel and AI Training

The value of data also comes from data itself. Companies that collect more data can build better targeting, attract more users, and gather even more data, creating a "data flywheel." For many companies, their data pile is now more valuable than the rest of their business.

  • Acquisitions for Data: Investors recognize that larger companies acquire smaller ones primarily for their data assets.
  • 23andMe: The bankrupt DNA testing company plans to sell its customers' genetic data as a strategic asset. Due to the nature of genealogy, even individuals who haven't used such services can have their information reverse-engineered through relatives in these databases.
  • Pokémon Go: Niantic's Pokémon Go was essentially a data collection tool, encouraging millions to point cameras at city landmarks to build a crowd-sourced 3D map. Niantic sold the game for $3.5 billion but kept the valuable mapping data, now used by an AI outfit to build geospatial models for robotics and autonomous systems.

The latest frontier in data mining is understanding how people think, to train AI models that are intended to automate human thought processes. This mirrors the automation of bank managers' roles, where individualized data made their discretion replicable and their jobs less important and cheaper to fill. The concern is that this could happen to every job.

AI companies are extremely data-hungry: - Anthropic: Bought millions of antique books, scanned them, and discarded the physical copies to feed its models, in addition to pirating books. - YouTube Videos: Subtitles from over 170,000 YouTube videos were scraped for training data by Apple, Nvidia, Anthropic, and others without creator consent. - Amazon/Twitch: Amazon scrapes its own users and Twitch streams for data. - Reddit: Google pays $60 million a year for Reddit comments, which are a prime source for AI training and frequently cited in AI search summaries.

The Irony of AI Data Collection

In a twist of irony, by feeding AI models, these companies have inadvertently undermined their own web traffic. When Google search returns an AI summary, users click through to actual websites only 8% of the time. - Stack Overflow: After selling data to train coding chatbots, new questions on Stack Overflow have plummeted to 2009 levels. - Business Insider: Google search traffic to Business Insider is down 85%. - USA Today: Traffic is down about 50%. - Reddit: Is renegotiating its $60 million deal with Google and considering blocking Google entirely.

While personal data is being exploited in new and creative ways, many of the companies facilitating this are also facing negative consequences from the very AI they helped train.

  Takeaways

  • Companies now treat personal data as raw material, aiming to generate roughly $2,000 per user each year, making data worth more than many individuals' incomes.
  • The practice dates back to early credit cards like Equifax’s “fact cards,” but digitalization and the FICO score standardized data, turning it into a scalable commercial asset.
  • Drastic drops in storage, collection, and analysis costs have enabled ubiquitous tracking from cars to DNA, turning previously impossible data volumes into profitable products.
  • Beyond advertising, data fuels employment verification services, dynamic pricing, investment models, political targeting, and AI training, expanding its value to wages, votes, and even crime statistics.
  • The data flywheel powers AI models, yet AI‑generated search summaries now divert traffic away from original sites, harming the very publishers that supplied the data.

Frequently Asked Questions

Why do analysts estimate companies must extract roughly $2,000 per user each year from personal data?

Analysts calculate that the high costs of collecting, storing, and processing massive data sets require a per‑user revenue of about $2,000 annually to be profitable. This amount covers infrastructure, data acquisition, AI training, and the premium advertisers pay for precise targeting, making the data operation financially viable.

What is “The Work Number” and how does it influence a person’s employment prospects?

“The Work Number” is Equifax’s employment‑verification database that stores detailed job histories, salaries, and re‑hire eligibility for millions of workers. Employers use it to screen candidates, often penalizing those with gaps or low prior wages, which can lower negotiating power and limit job opportunities.

Who is How Money Works on YouTube?

How Money Works is a YouTube channel that publishes videos on a range of topics. Browse more summaries from this channel below.

Does this page include the full transcript of the video?

Yes, the full transcript for this video is available on this page. Click 'Show transcript' in the sidebar to read it.

Why Data Collection Exploded

The shift from narrow financial applications to widespread data tracking is due to several factors:

Helpful resources related to this video

If you want to practice or explore the concepts discussed in the video, these commonly used tools may help.

Links may be affiliate links. We only include resources that are genuinely relevant to the topic.

Full transcript is not shown on this page

This page focuses on the summary and original notes. For full verification, refer to the original YouTube video.

PDF