Closeup detail of several Social Security Cards

In 2009 Study, CMU Team Used PSC Supercomputer and Public Online Data to Guess Social Security Numbers

Logo: 40 years of PSC

 

PSC40: Powering Discovery

2026 marks 40 years of PSC. As we continue on with cutting-edge innovation, we look back on four decades of history in computing, education, and groundbreaking research—and the people who made it happen.

In 2009, using PSC’s Pople system and routine information available online, a CMU team showed that you could guess at five or more of the nine digits in an individual person’s social security number up to 44 percent of the time. Many of their recommendations for fixing the SSN system remain unenacted. Their widely reported discovery showed that we are way too vulnerable to identity theft.

A supercomputer consisting of 7 towers
Pople was an SGI Altix 4700 shared-memory NUMA system comprising 192 blades, with a total of 768 cores.

WHY IT’S IMPORTANT

In 2023, identity theft cost U.S. adults $43 billion. That’s less than the $50 billion in estimated losses of 2007, but still a hefty wad of cash. And both numbers are an underestimate. Identity theft is notoriously under-reported. No wonder that, in 2025, we collectively spent $4.61 billion to protect ourselves from this crime.

While phishers and other scammers are happy to get a credit card or bank account number, the granddaddy of identity thefts revolves around finding a victim’s social security number. In 2008, CMU Professor Allessandro Acquisti and Ralph Gross, his postdoctoral researcher, wanted to follow up on some disturbing findings they’d made regarding the security of SSNs.

The problem, the CMU team realized, was that SSNs are assigned to people geographically. For older folks — those born before 1973 — this wasn’t as much of an issue. This was partly because many hadn’t gotten their numbers when they were born. Also, even in smaller states with higher population density, the lower frequency of assigning numbers in that era meant that time and place of birth offered fewer clues as to exactly what SSN would be assigned to a given person.

This began to change around the early ’70s, though, and went into overdrive after 1988. Particularly in small, densely populated states, the numbers began to be issued like clockwork. And that clockwork ticked out with disturbing regularity — and predictability. Worse, publicly available information online began to show exactly what “then” and “there” meant for each one of us.

The CMU scientists’ tool of choice to figure out what this information was inadvertently handing to bad actors was PSC’s Pople supercomputer, an SGI Altix 4700 shared-memory NUMA system.

HOW PSC HELPED

Acquisti’s and Gross’s initial work hinted that routine online information might be telling more than people realized about the probability of a given person being issued a specific SSN. They’d developed a statistical tool that could be used for exactly such a prediction, but it needed hefty computational power to predict more than the first few digits.

Pople, named for CMU’s Nobel Prize for Chemistry winner John Pople, was a monster at such work. Offering 768 computational cores that each had direct access to a then-enormous 1.5 terabytes of shared memory, it excelled at that kind of Concentration-like matching problem. In addition, to aid the CMU team’s work, PSC staff installed Octave, an open-source language that included powerful statistical functions, onto Pople.

The scientists began with the Social Security Administration’s Death Master File, or DMF. This publicly available database includes the SSNs, birth and death dates, and state of birth for every single deceased Social Security beneficiary. Since the SSNs belonged to dead people, the information in the DMF is harmless. But it could be used to test methods of guessing the numbers.

Leveraging the information in the DMF, Acquisti and Gross used Pople to guess at the SSNs of people who died between 1973 and 2003. Their guesses were way too accurate for comfort. If you were born between 1973 and 1988, the time and place of your birth allowed them to identify the first five digits of your SSN 7 percent of the time with just one guess. For those born after 1988, as the population got denser and the pace of issuing SSNs increased, that frequency rose to 44 percent.

Repeated guesses improved accuracy. With fewer than 1,000 guesses, they could determine all nine SSN digits for a person born after 1988 8.5 percent of the time. For people born in smaller states, it was even easier — 10 or fewer guesses provided the whole number 5 percent of the time.

The team’s report, in the Proceedings of the National Academy of Sciences USA in July 2009, made headlines across the nation. They’d demonstrated that then-current privacy procedures simply wouldn’t do. The practice of businesses using whole or partial SSNs for authentication, not strictly legal but common, was potentially ruinous. Most of all, Acquisti and Gross argued, the practice of assigning SSNs geographically was a ticking time bomb.

Today, few of their recommendations have been implemented. Assigning SSNs randomly as people are born would solve the issue going forward. But it wouldn’t help the more than 300 million people with already-assigned geographic numbers. Some kind of two-factor authentication might help, but again it hasn’t been enacted for SSNs.

The risk remains.