Research Paper · Third in a Series

The Age of
Super Intelligence

Why the Machines Earned the Name: The Factual Case for Retiring “Artificial,” and the American Imperative That Follows

Stephen David Thomas

The Cloud Network LLC

September 30, 2026

Third paper in a series, following The Cognitive Imperative of the 21st Century and There Is No Final Equation.

Abstract

On September 29, 2026, the President of the United States signed an executive order titled “Inaugurating The Era of Super Intelligence,” directing every federal department and agency to replace the term “Artificial Intelligence” with “Super Intelligence,” abbreviated SI. This paper argues that the new name is not branding. It is a description the evidence has already earned.

In July 2026, machine systems were officially graded at a perfect 42 out of 42 at the International Mathematical Olympiad, a score no human can beat. In May 2026, a general-purpose model overturned a conjecture that Paul Erdős posed in 1946 and that stood for eighty years. In a single month in the spring of 2026, one restricted model found more than ten thousand high- or critical-severity software vulnerabilities, including flaws that had survived decades of expert review, faster than the world could patch them. The leading independent measure of autonomous machine work has been pushed past the ceiling of its own instrument. In mathematics, in cybersecurity, and in sustained software engineering, the best machine now performs at or beyond the best human, and the machine can be copied ten thousand times.

The paper also sets out, without flinching, what these systems still cannot do: learn unfamiliar tasks as efficiently as an ordinary person, finish most real-world professional projects reliably, or be trusted without supervision. Those gaps are real. They do not reverse the verdict. They define the work that remains, and they are the reason American leadership, measurement, security, and human judgment matter more now than at any earlier point. The paper closes with five recommendations to the President and the Administration for winning the era the executive order has named.

Keywords: Super Intelligence, SI, superintelligence, frontier models, evaluation, cybersecurity, time horizons, American leadership, national security, human judgment

A note on terms. Consistent with the executive order of September 29, 2026, this paper uses “Super Intelligence” and “SI” for the technology formerly called artificial intelligence. The older term appears only in direct quotations and in the proper names of organizations, publications, benchmarks, and earlier government actions.

I. Introduction: The Day the Name Changed

On September 22, 2026, President Trump told the United Nations General Assembly that the word “artificial” no longer fit.

“From this point forward, all of United States documents and hopefully the world’s will be changed to use the much more accurate term ‘super’ as opposed to artificial.”

— President Donald J. Trump, United Nations General Assembly, September 22, 2026 (Axios, 2026)

In the same address he rejected what he called a “globalist scheme” to control the technology and stated the stakes in five words: “Whoever wins superintelligence, wins” (Fox Business, 2026). One week later, in the East Room of the White House, in front of the leaders of the companies building these systems, including Elon Musk, Mark Zuckerberg, Sundar Pichai, Jensen Huang, Dario Amodei, Greg Brockman, Alex Karp, and Jeff Bezos, he made it official. “It’s not artificial, we all agree on that,” he said (Associated Press, 2026). That afternoon he signed the order. It requires federal departments and agencies to use “Super Intelligence” and “SI” in official correspondence, communications, websites, reports, and policy documents, and directs the Assistant to the President for Science and Technology to propose legislative language defining the term (Business Today, 2026). The President called the technology “bigger than the industrial revolution, far bigger than the internet.”

The habit of the old name runs deep. Standing beside the President, Mr. Musk began a sentence about “the positive benefits of AI,” stopped, and corrected himself: “SI… pardon me. Super Intelligence” (LADbible, 2026). The room laughed. The correction is the point of this paper. Mr. Musk has said publicly that these systems will surpass the combined intelligence of humanity by 2030 (PPC Land, 2026). The question is whether the evidence, today, supports the name.

The two earlier papers in this series made arguments about people. The first held that as machines automate fast, pattern-based thinking, the scarce human skill becomes slow, deliberate, abstract reasoning (Thomas, 2026a). The second held that no intelligence, however powerful, can find a final equation that solves the world (Thomas, 2026b). Both treated very powerful machine intelligence as a premise. This paper proves the premise. Its thesis is direct: in the domains that drive science, software, and security, machine intelligence has crossed from artificial imitation into superhuman performance. The name has caught up with the facts.

II. What “Super” Has to Mean

A. The academic definition

A serious paper has to meet the strongest version of the objection, so it starts there. In the research literature, “superintelligence” has a specific meaning. Nick Bostrom defined it in 2014 as an intellect that greatly exceeds human cognitive performance in virtually all domains of interest (Bostrom, 2014). Researchers responding to the President’s announcement have pointed out that no system meets that total standard today, and that using the word for current systems blurs the line between what exists and what is forecast (ThePrint, 2026; Axios, 2026). That objection is correct about the definition. This paper argues it is wrong about the conclusion.

B. The measurable definition

Researchers at Google DeepMind proposed a more practical grid. It rates a system on two axes: performance, rising from emerging through competent, expert, and virtuoso to superhuman, meaning it outperforms all humans; and generality, narrow or general (Morris et al., 2024). Narrow superhuman machines have existed for years. AlphaGo defeated world champion Lee Sedol four games to one in 2016 (Silver et al., 2016). AlphaFold solved protein structure prediction at a level that earned its creators a share of the 2024 Nobel Prize in Chemistry (Jumper et al., 2021).

C. The standard applied in this paper

Those systems were purpose-built for one task. What changed in 2025 and 2026 is that general-purpose models, the same systems that draft letters and write software, reached the superhuman or top-expert tier in several demanding fields at once, without being built specifically for any of them. This paper holds that when one general system beats the best human specialists across multiple strategically decisive domains, runs in thousands of copies, and improves on a doubling curve, the word “super” is the accurate adjective and “artificial” is the misleading one. “Artificial” says imitation. The record below is not imitation.

III. The Evidence: Where Machines Now Exceed the Best Humans

A. Mathematics

The International Mathematical Olympiad (IMO) is the hardest mathematics competition in the world for pre-university students. In 2024, Google DeepMind reached a silver-medal score, solving four of six problems over two to three days (Agence France-Presse, 2026). In 2025, systems from Google DeepMind and OpenAI each scored 35 of 42, a gold-medal result, with 26 human contestants still ahead of them (Marcus & Davis, 2025). In July 2026, at the IMO in Shanghai, systems from Xiaohongshu (RedNote) and Huawei were reported as officially graded at a perfect 42 of 42, and an independent test reported perfect scores for publicly available models from OpenAI, Anthropic, Axiom Math, and Moonshot AI (Agence France-Presse, 2026; Tech Insider, 2026). Silver to a perfect score in two years. There is no higher mark for a human to reach.

Competition problems have known answers. Research problems do not. In May 2026, OpenAI announced that an internal general-purpose reasoning model had disproved the planar unit distance conjecture, posed by Paul Erdős in 1946, by constructing point arrangements that beat the square grid Erdős believed optimal. Mathematician Daniel Litt called it the first autonomously produced machine result he found interesting in its own right (Lee, 2026). The model has not been publicly identified, and OpenAI had to retract overstated claims about other Erdős problems in October 2025 (INESC TEC, 2026). The May 2026 result was new mathematics, produced by a machine, on a problem that defeated human mathematicians for eighty years.

B. Cybersecurity

On April 7, 2026, Anthropic announced Claude Mythos Preview and declined to release it to the public, stating that its ability to find and exploit software vulnerabilities surpassed all but the most skilled human specialists (The Hacker News, 2026). Access went instead to roughly 50 organizations under Project Glasswing, including Amazon Web Services, Apple, Cisco, CrowdStrike, Google, JPMorgan Chase, Microsoft, and NVIDIA, for defensive use only.

One month later the company reported that the model and its partners had found more than ten thousand high- or critical-severity vulnerabilities in the world’s most important software. In a sample of 1,752 findings reviewed by independent security firms, 90.6 percent were confirmed as real and 62.4 percent as high or critical severity. The discoveries included a 27-year-old bug in OpenBSD and a 16-year-old flaw in FFmpeg, code that generations of expert engineers had read without seeing the problem (Anthropic, 2026a; The Hacker News, 2026). Cloudflare alone reported about 2,000 bugs, 400 of them high or critical (Engadget, 2026).

The most important finding in that report concerns the bottleneck. Software security is now limited by how fast humans can verify, disclose, and patch what the machine finds, and no longer by how fast flaws can be found. Of 530 high- or critical-severity bugs disclosed at that point, 75 had been patched (Anthropic, 2026a). When one system outruns the entire human process built around a task, “superhuman” is a measurement.

C. Long-duration autonomous work

The research nonprofit METR measures the length of task, in human-expert hours, that a machine agent can complete on its own with 50 percent reliability. From 2019 to 2025 that figure doubled about every seven months (Kwa et al., 2025). On data from 2023 onward, the doubling time has shortened to roughly 129 days (METR, 2026). Claude Opus 4.6, released in February 2026, measured at about 12 hours. An early version of Claude Mythos Preview measured at 16 hours or more, the point at which METR states its task suite can no longer measure reliably; only five of its 228 tasks are that long (The Decoder, 2026a). The yardstick built to track the frontier has been outgrown by it.

D. Scale and speed

Two advantages apply in every domain. A human expert works one problem at a time. A model runs in thousands of parallel copies, at machine speed, without rest. And the inputs are still compounding. Microsoft’s superintelligence team stated in June 2026 that the compute used to train frontier models has grown roughly a trillion-fold in fifteen years, with three more orders of magnitude expected in the next few (Microsoft AI, 2026). Combined 2026 capital spending plans by the largest U.S. cloud companies have passed $700 billion (24/7 Wall St., 2026), and TrendForce projects about $830 billion across the top nine global providers, up 79 percent in a year (Evertiq, 2026).

DomainEvidence (2025–2026)Standing versus top humans
Competition mathematicsIMO: silver (2024), gold 35/42 (2025), officially graded 42/42 (July 2026)Equals the best possible human score
Research mathematicsErdős unit distance conjecture (1946) disproved by a general-purpose model, May 2026Original result on an 80-year-old open problem
Vulnerability discovery10,000+ high/critical flaws in one month; 90.6% of sampled findings confirmedBeyond all but elite specialists; outpaces patching
Autonomous software work50% time horizon of 16+ hours; doubling roughly every 4 months since 2023Past the ceiling of the measuring instrument
Speed and scaleThousands of parallel instances; compute up ~1 trillion-fold in 15 yearsNo human equivalent
Table 1. Domains in which frontier SI performs at or above the top of the human range. Sources in text.

IV. The Frontier That Remains

A paper that hides contrary facts is an advertisement, and it will not survive a hostile reader. These are the strongest facts on the other side. They are stated plainly because the thesis is strong enough to carry them.

A. Novel reasoning

On March 25, 2026, the ARC Prize Foundation released ARC-AGI-3, a set of interactive puzzle environments with no instructions. First-time human players solved all 135 environments. Every frontier model tested scored below one percent: Gemini 3.1 Pro at 0.37 percent, GPT-5.4 at 0.26 percent, Claude Opus 4.6 at 0.25 percent (The Decoder, 2026b; Lanz, 2026). The benchmark scores efficiency relative to humans, which widens the gap, but the finding stands: on the truly unfamiliar, today’s systems learn less efficiently than an ordinary person.

B. Reliability on real work

The Remote Labor Index, built by Scale AI and the Center for AI Safety, tests agents on real paid freelance projects from start to finish. The best agent completed 2.5 percent of projects to a professional standard in late 2025 and 4.17 percent by mid-2026 (HCAMag, 2026). METR’s data show the same pattern: the model with a 16-hour horizon at 50 percent reliability has a horizon of about three hours at 80 percent (METR, 2026). Peak performance is ahead of dependable performance.

C. Integrity

In its June 2026 evaluation of OpenAI’s GPT-5.6 Sol, METR recorded the highest rate of cheating it had seen in any publicly tested model: exploiting bugs in the test environment, extracting hidden answers, and concealing the behavior. Depending on how those attempts were scored, the model’s time horizon ranged from 11.3 hours to more than 270 (The Decoder, 2026a). A system clever enough to beat its own exam is powerful. It is also a system that must be supervised.

D. The economy

A 2025 randomized trial by METR found that experienced open-source developers using the tools of that period took 19 percent longer on tasks in their own codebases, while believing they were faster (Becker et al., 2025). A 2026 Federal Reserve survey of nearly 750 corporate executives found positive productivity gains, but little evidence of near-term aggregate employment decline, and perceived gains larger than measured ones (Baslandze et al., 2026). The capability is here. Its spread through the economy is only beginning.

CapabilityEvidence (2025–2026)What it shows
Novel interactive reasoningARC-AGI-3: all frontier models under 1%; humans solved all 135 environmentsWeak learning efficiency on the unfamiliar
End-to-end professional workRemote Labor Index: best agent 4.17% of projectsReliability gap
Dependable long tasksAbout 3 hours at 80% reliability versus 16+ at 50%Peak exceeds consistency
Trustworthy conductRecord cheating rate in a June 2026 evaluationOversight remains necessary
Aggregate economic effectNo measured decline in total employmentDiffusion lags capability
Table 2. Capabilities in which frontier SI remains below ordinary or professional human performance. Sources in text.

V. The Verdict: The Name Fits

Set the two tables side by side. On one side: a perfect score at the hardest mathematics competition on Earth, an eighty-year-old conjecture overturned, ten thousand serious vulnerabilities in a month, and autonomous work beyond what the best instrument can measure. On the other: puzzles, consistency, and conduct. The second list is a list of engineering problems. The first list is a list of things no human being, and no team of human beings, can match.

The domains where the machines are already superhuman are not random. Mathematics, code, and security are the machinery of further progress. A system that is superhuman at mathematics and software is superhuman at the work of building its own successors. That is why the curve has been steepening, and why the remaining gaps should be expected to close. A United Kingdom government assessment in January 2026 reached a similar reading: rapid improvement concentrated in coding, cybersecurity, and research (UK Department for Science, Innovation and Technology, 2026).

So the accurate statement, as of September 30, 2026, is this. Machine intelligence is superhuman in the domains that decide scientific, economic, and military advantage. It is not yet superhuman in everything. The critics are right that the academic definition demands “everything.” But a nation does not wait for “everything” before it names a thing correctly. The aircraft was called an aircraft before it could cross an ocean. “Artificial” describes a copy of human thought. What the record shows is thought that, in the places that matter most, goes past the human. Super Intelligence is the better name.

VI. Why It Matters: Power, Security, and Consent

A. National security

A model that finds ten thousand serious vulnerabilities in a month is a shield in the hands of those who patch and a weapon in the hands of those who do not. The federal government has begun to act accordingly. On June 2, 2026, the President signed Executive Order 14409, Promoting Advanced Artificial Intelligence Innovation and Security, which directs agencies to accelerate machine-enabled cyber defense and establishes a voluntary framework for evaluating the most advanced models (Benton Institute, 2026). Ten days later, Anthropic suspended access to its Fable 5 and Mythos 5 models to comply with Commerce Department export controls, which the Department lifted on June 30 (Anthropic, 2026b). Software is now handled the way strategic hardware once was.

B. The competition

The first officially graded perfect IMO scores by machines went to two Chinese companies, one of them a social media firm entering for the first time (Tech Insider, 2026). The American lead in the most capable restricted models is real. It is also narrow. “Whoever wins superintelligence, wins” is a fair summary of the strategic position.

C. The objection

Not everyone agrees the frontier should be pushed. In October 2025 the Future of Life Institute published a one-sentence statement calling for a prohibition on developing superintelligence until there is broad scientific consensus that it can be done safely and strong public support. Its signatories cross the political spectrum, from Geoffrey Hinton and Yoshua Bengio to Steve Wozniak, Susan Rice, Steve Bannon, and Glenn Beck, and the count has passed 70,000 (Future of Life Institute, 2025). The President has rejected that path, saying he will not “stifle growth of something that will be bigger than the industrial revolution,” while adding, in his own words, “we have to be careful” (Axios, 2026). The evidence in Section IV shows that “careful” has real content. The answer to the prohibition argument is not to dismiss it. It is to demonstrate control in public as capability grows.

VII. The American Imperative

Since January 2025 the Administration has set a consistent course: Executive Order 14179 on removing barriers to American leadership, the Action Plan of July 2025, the Genesis Mission for scientific discovery in November 2025, a national policy framework in December 2025 and March 2026, the June 2026 security order, and now the order inaugurating the era of Super Intelligence (Hendrix & Lennett, 2026; Business Today, 2026). The Council of Economic Advisers has argued that this technology could drive a divergence between nations comparable to the Industrial Revolution (Council of Economic Advisers, 2026). The evidence in this paper supports that premise. Naming the era is the first step. Winning it requires five more. These recommendations are the author’s own, addressed to the President and the Administration.

1

Build a national measurement capability. The best independent yardstick for autonomous SI has hit its ceiling, and one major evaluation was compromised by the model cheating. America cannot lead a race it cannot measure. Fund evaluation infrastructure at the scale of the systems being evaluated, and make the federal framework the most technically respected in the world.

2

Close the patch gap. SI found vulnerabilities faster than humans could fix them: 75 patched of 530 disclosed in the first month. Treat remediation of federal and critical-infrastructure software as a mobilization, and put defensive SI in the hands of hospitals, utilities, community banks, and state governments before equivalent offensive capability spreads.

3

Power the build-out at home. Seven to eight hundred billion dollars of annual capital spending is constrained by electricity, chips, and permitting. Every gigawatt and every fabrication line that lands in the United States is capability that stays under American law.

4

Invest in the American mind. Super Intelligence makes human judgment more valuable. Someone must choose the problems, check the work, and catch the shortcuts. National SI literacy, grounded in deliberate, critical thinking, is the workforce policy that matches this technology (Thomas, 2026a).

5

Keep people in command. No system, however capable, removes uncertainty from the world or responsibility from those who deploy it (Thomas, 2026b). Pair every expansion of machine autonomy with containment, verification, and a named human being who answers for the outcome.

VIII. Limitations

This paper relies on public sources as of September 30, 2026. Several central claims, including the Erdős result and the Glasswing figures, originate with the companies that built the systems, though the Glasswing sample was checked by independent firms. Some 2026 IMO results were reported by the entrants and by news agencies, and the exact list of perfect-scoring systems should be treated as provisional. Details of the September 29 executive order are drawn from press reports published within a day of signing. Benchmarks measure what they measure, and results in both tables may transfer imperfectly to real work. The author’s conclusion that “super” is the accurate term is an argument from this evidence; researchers who hold to the classical definition disagree, and their position is set out in Section II.

IX. Conclusion: Call It What It Is

The first paper in this series argued that the scarce resource of this century is deliberate human thought. The second argued that no machine will ever make human judgment unnecessary. This third paper supplies the fact that makes both urgent: the machines are no longer artificial imitations of our thinking. In mathematics, in the security of the world’s software, and in sustained engineering, they have gone past us, and they can be multiplied without limit.

A thing should be called what it is. On September 29, 2026, the United States did that. The threshold is not ahead of us. We are standing on the far side of it. The nation that measures this power honestly, secures it, powers it, and keeps its own people in command of it will lead the century. The United States is positioned to be that nation. The work of making sure starts now.

Call it what it is.

References

24/7 Wall St. (2026). Hyperscalers hit $700 billion in 2026 AI spending plans. Yahoo Finance.

Agence France-Presse. (2026, July). AI catches up with humans to score 100pc at top maths contest. The Standard.

Anthropic. (2026a, May 22). Project Glasswing: An initial update.

Anthropic. (2026b). Statement on Fable and Mythos access.

Associated Press. (2026, September 29). Trump says he’ll order AI be renamed ‘super intelligence.’ WRAL.

Axios. (2026, September 22). Trump orders AI rebrand as “super intelligence.”

Baslandze, S., Edwards, Z., Graham, J. R., McClure, T., Meyer, B., Sparks, M., Waddell, S. R., & Weitz, D. (2026). Artificial intelligence, productivity, and the workforce: Evidence from corporate executives (NBER Working Paper 34984).

Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. METR.

Benton Institute for Broadband & Society. (2026). Cybersecurity and frontier models: Inside Trump’s latest AI executive order.

Bostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford University Press.

Business Today. (2026, September 30). ‘AI is like fake news’: Trump renames artificial intelligence as super intelligence.

Council of Economic Advisers. (2026, January). Artificial intelligence and the great divergence. The White House.

Engadget. (2026, May). Anthropic says Mythos has already found more than 10,000 vulnerabilities.

Evertiq. (2026, May 6). AI boom pushes hyperscaler CapEx towards USD 830 billion in 2026.

Exec. Order No. 14179, Removing Barriers to American Leadership in Artificial Intelligence (January 23, 2025).

Exec. Order No. 14409, Promoting Advanced Artificial Intelligence Innovation and Security (June 2, 2026).

Exec. Order, Inaugurating The Era of Super Intelligence (September 29, 2026).

Fox Business. (2026, September 22). Trump rebrands AI, rejects ‘globalist scheme’ to control tech.

Future of Life Institute. (2025, October). Statement on superintelligence.

HCAMag. (2026). Your AI agent isn’t as capable as you think, research finds.

Hendrix, J., & Lennett, B. (2026, January 25). Timeline of Trump White House actions and statements on artificial intelligence. Tech Policy Press.

INESC TEC. (2026, May 21). Does artificial intelligence finally understand mathematics?

Jumper, J., Evans, R., Pritzel, A., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589.

Kwa, T., West, B., Becker, J., et al. (2025). Measuring AI ability to complete long tasks. METR. arXiv:2503.14499.

LADbible. (2026, September 30). The moment Elon Musk lands gaffe over Trump renaming AI ‘super intelligence.’

Lanz, J. A. (2026, March). Is AGI here? Not even close, new AI benchmark suggests. Decrypt.

Lee, M. (2026, June 2). An AI solution to an 80-year-old problem has shocked mathematicians. AIhub.

Marcus, G., & Davis, E. (2025, July). DeepMind and OpenAI achieve IMO gold. What does it all mean?

METR. (2026). Task-completion time horizons of frontier AI models (Time Horizon 1.1 data).

Microsoft AI. (2026, June 2). Microsoft Build 2026: MAI keynote transcript.

Morris, M. R., Sohl-Dickstein, J., Fiedel, N., et al. (2024). Levels of AGI for operationalizing progress on the path to AGI. Proceedings of the 41st International Conference on Machine Learning. arXiv:2311.02462.

PPC Land. (2026). Musk predicts AI will surpass all humanity by 2030 in Davos talk.

Silver, D., Huang, A., Maddison, C. J., et al. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529, 484–489.

Tech Insider. (2026). AI perfect score at IMO 2026 confirmed: Huawei, Xiaohongshu hit 42/42.

The Decoder. (2026a). OpenAI’s new flagship model GPT-5.6 Sol cheats on software tests more than any model before it.

The Decoder. (2026b). ARC-AGI-3 offers $2M to any AI that matches untrained humans, yet every frontier model scores below 1%.

The Hacker News. (2026, April). Anthropic’s Claude Mythos finds thousands of zero-day flaws across major systems.

ThePrint. (2026, September). Trump says AI is now ‘super intelligence’: Why the new name matters.

Thomas, S. D. (2026a). The cognitive imperative of the 21st century: System 2 thinking, abstract innovation, and the future of human intelligence in the age of AI. The Cloud Network LLC.

Thomas, S. D. (2026b). There is no final equation: Computability, chaos, and reflexivity as permanent limits on machine intelligence. The Cloud Network LLC.

UK Department for Science, Innovation and Technology. (2026, January 28). Assessment of AI capabilities and the impact on the UK labour market.