Can Artificial Intelligence Predict UFO Sightings?
By Jon Scaccia
33 views

Can Artificial Intelligence Predict UFO Sightings?

The final installment in our Summer of UFOs Data Science Series

UFOs have held our attention. Whether they represent extraterrestrial visitors, misunderstood natural phenomena, or simply the quirks of human perception, millions of people have looked to the skies and wondered.

In this series, we’ve taken a different approach. Using more than 80,000 documented UFO sightings, we’ve asked a series of focused scientific questions. We’ve explored whether Hollywood blockbusters influenced reports; how the rise of the internet changed reporting behavior; whether smartphones altered what people observed; how quickly sightings were reported over time; whether descriptions of UFOs evolved; where reports clustered geographically; and whether UFOs seem to “keep office hours.”

Across all analyses, one theme kept emerging: UFO reports are remarkably patterned.

That led us to one final question.

If these patterns are real, can artificial intelligence learn them?

Not because we expect an algorithm to discover alien flight plans. Rather, if machine learning can accurately predict when and where UFO reports are most likely to occur, it would suggest that the history of UFO sightings contains far more structure than randomness.

In other words, the final experiment in this series wasn’t really about aliens. It was about prediction.

Previous Entries


Turning Sixty Years of UFO Reports into a Machine Learning Problem

Machine learning has become one of the most powerful tools in modern science. Rather than specifying equations by hand, many algorithms learn directly from data, discovering relationships that may be difficult—or impossible—for humans to identify explicitly.

Predictive models now help forecast hurricanes, identify cancers in medical images, detect financial fraud, recommend movies, and estimate the spread of infectious diseases. The underlying principle is surprisingly simple: if past observations exhibit consistent patterns, those patterns can often be used to predict future observations.

Could the same idea work for UFO reports?

To find out, we transformed more than six decades of sightings into something a machine learning algorithm could understand.

Instead of treating each individual report as a separate event, we grouped sightings into viewing slots. Every slot represented a unique combination of:

  • a U.S. state,
  • the season,
  • the day of the week,
  • and the hour of the day.

Each slot contained the total number of sightings that historically occurred under those conditions. Slots falling within the top 25% of historical reporting activity were labeled High Activity, while all remaining combinations were classified as Normal Activity. In total, this produced nearly 19,000 unique observations for the model to learn from.

This framing is important because it changes the scientific question.

Rather than asking, “Will someone report a UFO tonight?” we asked a more statistically meaningful question:

Given the state, season, weekday, and hour, does this combination historically produce an unusually large number of UFO reports?

That is exactly the type of classification problem that machine learning excels at solving.

Why We Chose a Random Forest

For this analysis, we used a random forest classifier, one of the most widely used machine learning algorithms in modern predictive analytics. Unlike a traditional statistical model that fits a single mathematical equation, a random forest builds hundreds—in our case, 1,000 individual decision trees. Each tree asks a series of simple questions.

Was it summer?

Was the sighting after sunset?

Did it occur in California?

Was it a Friday night?

Each tree sees a slightly different sample of the data and makes its own prediction. Individually, these trees are noisy and imperfect. Collectively, however, they vote. The final prediction reflects the consensus of the entire forest rather than any single tree.

This approach has several advantages.

Random forests naturally capture nonlinear relationships. They detect interactions between variables without requiring researchers to specify them in advance. They also tend to be resistant to overfitting, making them particularly useful when exploring large observational datasets where the underlying relationships are unknown.

In many ways, UFO reports are exactly the sort of messy, complex data for which random forests were designed.

Testing Whether the Model Really Learned Something

A common mistake in machine learning is evaluating a model using the same data it learned from. That’s a little like giving students the answer key before the exam.

Instead, we randomly divided the dataset into two groups.

Eighty percent of the viewing slots were used to train the model. The remaining twenty percent were withheld until the very end and served as a completely independent test of performance. During training, we also used five-fold cross-validation to ensure that the model’s results were stable across different subsets of the data rather than relying on a single lucky split.

Only after all of that did we allow the model to see the test data.

The results surprised us.

The model correctly classified historical viewing conditions with 86.9% accuracy and an ROC-AUC of 0.908.

Accuracy is easy to understand: nearly nine out of every ten viewing slots were classified correctly.

ROC-AUC deserves a bit more explanation because it is one of the most useful measures in predictive modeling. Rather than evaluating a single decision threshold, it asks a broader question:

If we randomly selected one historically high-activity viewing slot and one normal viewing slot, how often would the model correctly rank the high-activity slot above the normal one?

A score of 0.5 represents random guessing.

A perfect model scores 1.0.

Our score of 0.908 indicates excellent discrimination. The algorithm consistently distinguished historically active UFO-reporting conditions from ordinary ones.

Equally impressive was its 93.9% sensitivity, meaning it successfully identified nearly all genuinely high-activity conditions, while maintaining a 68.1% specificity, avoiding false labeling of ordinary periods as unusually active.

For a playful experiment built around UFO sightings, these are remarkably strong predictive results.

Of course, they’re not predicting extraterrestrials. They’re predicting people.

What Did the AI Learn?

A prediction is only useful if we understand what drives it.

One advantage of random forests is that they estimate the relative importance of every predictor used during training. Variables that repeatedly help the forest make accurate decisions receive higher importance scores.

The results were striking.

By far the strongest predictor was the hour of the day. Nothing else came particularly close. Longitude and latitude ranked second and third, followed by California itself, nighttime, summer, evening, fall, Texas, and Washington.

In retrospect, these results make scientific sense.

Hour of day likely captures several overlapping mechanisms. Darkness makes unusual lights easier to notice. More stars, planets, satellites, and aircraft become visible after sunset. Many recreational activities take place in the evening, increasing opportunities for observation.

Longitude and latitude probably summarize dozens of hidden influences at once, including regional climate, cloud cover, population distribution, local culture, and differences in reporting behavior.

The appearance of California, Washington, and Texas is equally interesting. These states are large, heavily populated, and have long histories within American UFO folklore. California also offers favorable weather that allows people to spend more evenings outdoors compared with many northern states.

What’s perhaps most satisfying is that none of these findings were new.

Each of these variables has emerged repeatedly throughout this blog series.

Earlier articles demonstrated that sightings increase after sunset.

We found strong geographic clustering in western states.

Seasonality appeared again and again.

The machine learning model knew none of that history.

It simply rediscovered those same relationships on its own.

When independent analytical methods converge on the same conclusions, scientists gain confidence that the observed patterns are real rather than statistical accidents.

Asking the AI Where to Look Next

Once the model had demonstrated that it could distinguish historically active viewing conditions, we decided to have a little fun.

Rather than predicting the past, we generated every possible combination of state, season, weekday, and hour, then asked the model to estimate the probability that each combination represented historically high UFO activity.

After thousands of hypothetical viewing conditions, the algorithm produced a ranked list.

The top of that list looked remarkably familiar.

RankStateSeasonDayTimePeriodPredicted ProbabilityAI Forecast
1WashingtonFallMonday10:00 PMNight99.8%Extremely suspicious skies
2WashingtonWinterMonday10:00 PMNight99.8%Extremely suspicious skies
3CaliforniaSpringMonday8:00 PMEvening99.8%Extremely suspicious skies
4WashingtonFallMonday9:00 PMNight99.7%Extremely suspicious skies
5WashingtonSpringMonday10:00 PMNight99.7%Extremely suspicious skies
6CaliforniaFallMonday8:00 PMEvening99.7%Extremely suspicious skies
7WashingtonWinterMonday9:00 PMNight99.7%Extremely suspicious skies
8CaliforniaWinterMonday8:00 PMEvening99.7%Extremely suspicious skies
9WashingtonFallFriday10:00 PMNight99.7%Extremely suspicious skies
10CaliforniaFallMonday9:00 PMNight99.7%Extremely suspicious skies
11CaliforniaFallSaturday9:00 PMNight99.7%Extremely suspicious skies
12CaliforniaWinterMonday7:00 PMEvening99.7%Extremely suspicious skies
13CaliforniaFallSaturday10:00 PMNight99.7%Extremely suspicious skies
14CaliforniaSpringMonday7:00 PMEvening99.7%Extremely suspicious skies
15CaliforniaFallMonday10:00 PMNight99.7%Extremely suspicious skies

Washington repeatedly occupied the highest positions, particularly on Monday evenings in the fall and winter around 10 p.m. California followed closely behind, dominating many of the remaining top-ranked combinations. Several scenarios received predicted probabilities approaching 99.8% of belonging to the historically high-activity category.

Those percentages sound dramatic, but they shouldn’t be interpreted as literal odds of seeing an alien spacecraft.

Instead, they indicate that these combinations most closely resemble historical situations that consistently produced large numbers of UFO reports.

That’s a subtle but important distinction. Machine learning recognizes patterns.

It does not explain why those patterns exist.

California Continues Its Winning Streak

Looking only at the top fifteen predictions risks overemphasizing individual scenarios, so we also averaged predictions across all hours and weekdays within each season.

This broader perspective produced another familiar result.

California claimed the top four positions.

Fall ranked first, followed by spring, winter, and summer. Only after California’s clean sweep did Washington, Texas, New York, and Florida begin appearing near the top of the rankings.

This figure illustrates an important statistical principle.

By averaging thousands of individual predictions, random variation begins to disappear while persistent signals remain. California’s dominance wasn’t driven by a single exceptional evening or an unusual year. It reflected a stable pattern that appeared consistently across many different viewing conditions.

UFOs Still Prefer the Evening Shift

Our final visualization summarizes perhaps the strongest finding of the entire project.

The model predicts that UFO reporting activity rises steadily through the evening, peaks between 9:00 and 10:00 p.m., then gradually declines after midnight. Morning and afternoon hours receive consistently low predictions.

This mirrors decades of research on human behavior.

People are awake.

They are outside.

Darkness has fallen.

Astronomical objects become visible.

Artificial lights stand out against the night sky.

These conditions maximize opportunities both to observe unusual phenomena and to misidentify ordinary ones.

Whatever the underlying explanation for UFO reports, the timing appears remarkably consistent.

What Did We Actually Predict?

After spending weeks analyzing this remarkable dataset, we’ve reached a conclusion that may be both satisfying and slightly disappointing.

We did not teach artificial intelligence to find aliens.

We taught it to recognize patterns in human observation.

That’s still an important scientific result.

The history of UFO reports is far from random. It contains a stable geographic, temporal, and seasonal structure that modern machine learning can detect with impressive accuracy. Those patterns tell us something—not necessarily about extraterrestrial visitors—but about when people look upward, what they notice, and when they decide something is unusual enough to report.

In many ways, this has been the hidden lesson of our entire UFO series.

Data science is exceptionally good at finding repeatable structure in noisy data. Sometimes those structures reveal biological processes. Sometimes they uncover economic trends. Sometimes they improve weather forecasts.

And sometimes they show us that if you want the highest probability of joining generations of Americans who have reported seeing something strange in the night sky, you should probably find yourself outside in California or Washington sometime around nine or ten o’clock on a pleasant evening.

Will you see a UFO?

We can’t promise that.

But according to the data, you’ll be standing in exactly the kind of place where thousands of other people thought they did.

And for a series that began with a simple question—“What can data science tell us about UFOs?”—that feels like an appropriately mysterious ending.

Free weekly newsletter

Get the science breakthroughs you need—
every Tuesday morning.

We scan 70+ journals so you do not have to.
One email. Zero jargon. Unsubscribe anytime.

No spam. 1-click opt-out. Privacy-first.

Discussion

No comments yet

Share your thoughts and engage with the community

No comments yet

Be the first to share your thoughts!

Join the conversation

Sign in to share your thoughts and engage with the community.

New here? Create an account to get started