---
title: "The new model 4.6 of Claude is here! More jobs are gone: Wall Street finance, compilers, security white hats, PPT... all have fallen"
type: "News"
locale: "en"
url: "https://longbridge.com/en/news/275069265.md"
description: "The release of Claude Opus 4.6 new model caused a 10% plunge in the stock price of Wall Street financial data service provider FactSet, with S&P Global, Moody's, and Nasdaq also experiencing declines. Investor panic over AI disruption has intensified. The new model excels in areas such as financial analysis, research, and Office applications, with officials stating it surpasses OpenAI GPT-5.2 in the GDPval-AA assessment. The pricing for Opus 4.6 remains unchanged, with an input/output price of $5/$25 per million tokens"
datetime: "2026-02-06T03:59:29.000Z"
locales:
  - [zh-CN](https://longbridge.com/zh-CN/news/275069265.md)
  - [en](https://longbridge.com/en/news/275069265.md)
  - [zh-HK](https://longbridge.com/zh-HK/news/275069265.md)
generator: "portal-rs"
---

# The new model 4.6 of Claude is here! More jobs are gone: Wall Street finance, compilers, security white hats, PPT... all have fallen

As soon as I opened my eyes, Anthropic launched a new model, let Claude Opus 4.6 come to wish you a Happy New Year!

As soon as the news broke, financial data service provider FactSet plummeted 10% during trading, while S&P Global, Moody's, and Nasdaq all saw declines, with major indices diving across the board.

**This is already the second time this week that Anthropic has stirred the market**.

A few days ago, a plugin for automated legal work quietly went live under its banner, directly triggering a plunge in trillion-dollar software stocks.

Investor panic is focused on one question: who can guarantee they won't be disrupted by AI in a few years? If not, they sell off.

Little did we expect that today's Anthropic would be even more ruthless.

Before today, everyone's impression of Claude was that its programming capabilities were intermittently strong.

Claude Opus 4.6 sneered and punched through this impression: I'm strong in many more areas!

At least according to official statements, in financial analysis, research, and the Office suite, Claude Opus 4.6 can perform very well.

The official website states:

> In GDPval-AA (a performance metric assessing economic value knowledge work tasks in finance, law, and other fields), Opus 4.6 outperformed the industry's next best model, OpenAI GPT-5.2, by 144 Elo!

\_ (This means that Claude Opus 4.6 scores higher than GPT-5.2 in about 70% of cases in this assessment, with 50% indicating equivalent scores)\_

Of course, it still leads the way in programming.

It achieved the highest score in the Agent programming assessment Terminal-Bench 2.0 and outperformed all other cutting-edge models in the "final human exam."

The good news is that the quantity increases without a price increase; the pricing for Opus 4.6 remains at the original standard: **$5 per million tokens for input/output, $25** *(For ease of reading, hereinafter referred to as the new model Opus 4.6)*

## **Returning to the Peak with 1M Context and Adaptive Thinking**

The most intuitive advancement of Opus 4.6 is the **introduction of a 1M Token super large context**, which is the first time Claude has introduced this length of context window in an Opus-level model.

This greatly improves Opus 4.6's performance in handling long texts, alleviating the issue of "context decay."

In the MRCR v2 8-needle 1M benchmark test—finding a needle in a haystack—Opus 4.6 scored 76%, while Claude Sonnet 4.5 only scored 18.5%.

The accompanying result is an enhancement in search capabilities.

In the BrowseComp evaluation *(assessing the ability to retrieve hard-to-find information online)*, Opus 4.6 ranked first in the industry, with the best performance in deep multi-step agent-style searches, accurately pinpointing key information scattered throughout long documents.

Opus 4.6 also introduces the Adaptive Thinking feature.

Previously, developers using Claude models could only choose one of two options: to enable or disable the expanded thinking mode.

Now, Claude can autonomously determine when deep reasoning is necessary.

*(To be honest, this step is slower than ChatGPT, so please speed up the introduction of such good features next time.)*

The accompanying effort parameter provides four options—low, medium, high, max—with the default set to high, allowing manual adjustment to lower levels in cases of excessive model thinking.

Another practical feature is Context Compaction.

When the conversation approaches the upper limit of the context window, it automatically summarizes and replaces old content, making long conversations and agent tasks easier.

## **Core Scenarios like Coding, Knowledge Work, Search, and Reasoning Have Exploded**

The official blog shows that as soon as Opus 4.6 was released, almost no model could compete with it.

**In core scenarios such as coding, knowledge work, search, and reasoning, Opus 4.6 has made significant breakthroughs.** Multiple evaluation scores surpass previous generations and industry competitors, be like:

After getting a general impression, let's break it down one by one.

**First is programming capability.**

Opus 4.6 achieved the highest score in Terminal-Bench 2.0.

From the actual capabilities behind the scores, Opus 4.6 can plan tasks more meticulously, run stably in large codebases, and improve code review and debugging accuracy.

Moreover, it can autonomously identify its own errors.

Another point is that Opus 4.6 supports multi-language coding and can handle cross-language software engineering issues.

It can complete the migration of millions of lines of code like a senior engineer, and it takes half the time.

As I write this, I can't help but wonder:

Are engineers happy enough that their hair won't fall out, or will it fall out even faster... *（陷入沉思.jpg）*

**Secondly, Opus 4.6 is also actively invading traditional office territory.**

This time it has made a strong move against the Office trio.

-   It can directly ingest messy unstructured data in Excel, autonomously infer a reasonable table structure, and handle multiple complex steps in one operation;
-   It can remember your company's PPT templates, including fonts and layout styles, ensuring that the generated PPT has no AI flavor, making the boss think you stayed up late to put it together.

**In a coworking environment, Opus 4.6 can autonomously multitask on behalf of the user, running financial analysis while organizing research results into documents.**

It feels like Anthropic wants to pull Claude out of the chat box into more spaces?

**Third, let's talk about its progress in reasoning ability.**

First, a summary:

> Opus 4.6 is stronger in cross-domain reasoning.

In the multidisciplinary complex reasoning test "The Last Exam for Humans," Opus outperformed all leading models.

In the legal field, Opus 4.6 scored 90.2% on the BigLaw Bench, where 40% is the full score.

In economic value-oriented task evaluations like GDPval-AA in finance and law, Opus 4.6 surpassed the "industry competitor" OpenAI GPT-5.2 with a score of 144 Elo Whether it's complex legal or financial expertise or intricate academic research, its reasoning and understanding depth have reached the peak of current frontier models.

What's rare is that **this leap in intelligence has not come at the cost of safety**.

In the automated behavior auditing that Anthropic values most, Opus 4.6 has a **very high alignment level, while negative behaviors such as deception and flattery are extremely low**.

Opus 4.6 even addresses the "over-rejection" problem that is currently a common headache in the AI community—

When faced with normal, harmless requests, it exhibits that rigid refusal less than any previous model.

Currently, Opus 4.6 has been launched on the official website, API, and all major cloud platforms.

No increase in price, Opus 4.6's pricing remains at the original standard: **$5 per million tokens for input/output, $25**.

However, in the 10M token context testing version, there will be additional charges if the prompt exceeds 200k tokens.

Key Point!

To use Opus 4.6, you need to explicitly specify the model identifier "Claude-opus-4-6" when calling the API.

## **More Jobs Gone**

**16 Agents wrote a C compiler in two weeks, running Doom**

A core capability upgrade brought by Opus 4.6 is Agent Teams, which allows multiple Claude instances to collaborate in parallel without real-time human supervision.

Nicholas Carlini, a researcher from Anthropic's safety team, conducted a stress test: he had 16 Agents write a C compiler capable of compiling the Linux kernel from scratch using Rust.

In two weeks, nearly 2000 Claude Code sessions consumed 2 billion input tokens and 140 million output tokens, with a total cost of less than $20,000.

The final output was a 100,000-line compiler that could compile Linux 6.9 on x86, ARM, and RISC-V architectures and could run Doom.

This parallel mechanism allows each Agent to run in an independent Docker container, sharing a single git repository.

To prevent multiple Agents from colliding and rushing to solve the same problem, the system employed a simple locking mechanism.

Agents "claim" tasks by writing files to the current\_tasks/ directory, and git's synchronization mechanism automatically handles conflicts. There is no dedicated communication protocol between Agents, nor is there orchestration of Agents; each Claude decides what to do next on its own Carlini wrote on the blog:

"When the Agent started compiling the Linux kernel, it got stuck at one point because it was a massive monolithic task, and 16 Agents all hit the same bug and overlapped with each other."

The solution was to introduce GCC as an "oracle" control group, allowing each Agent to compile a random subset of the kernel, thereby using binary search to locate the problematic files, which truly unleashed the parallel capability.

**500 Zero-Day Vulnerabilities, Digging Right Out of the Box**

Opus 4.6's performance in the cybersecurity field surprised even Anthropic itself.

In pre-release testing, Anthropic's cutting-edge red team threw Opus 4.6 into a sandbox environment, providing it with Python and conventional vulnerability analysis tools (such as fuzzers and debuggers), without any specific instructions or domain knowledge, letting it find vulnerabilities in open-source code on its own.

As a result, it **unearthed over 500 previously unknown high-risk zero-day vulnerabilities**.

Each one was verified by members of the Anthropic team or external security researchers.

Specific cases include:

-   A crash-causing vulnerability discovered in GhostScript (a common tool for processing PDF and PostScript files), which was found after traditional fuzzing and manual analysis failed to identify the issue; Claude dug it out by reviewing the project's git commit history.
-   Buffer overflow vulnerabilities found in OpenSC (a tool for handling smart card data) and CGIF (a tool for processing GIF files); in the CGIF case, Claude even proactively wrote a PoC (proof of concept) to demonstrate the existence of the vulnerability.

Logan Graham, head of Anthropic's cutting-edge red team, said he wouldn't be surprised if this became one of the main methods for auditing the security of open-source software in the future.

However, Anthropic also acknowledges that this capability could be misused.

To address this, the team has added six new cybersecurity detection mechanisms, and a real-time interception system may be launched in the future to block malicious traffic.

## **One More Thing**

The official website shows that Anthropic is now "building Claude with Claude."

Their engineers use Claude Code to write code every day, and each new model is first tested in their own working environment.

 Risk Warning and Disclaimer

The market has risks, and investment should be cautious. This article does not constitute personal investment advice and does not take into account the specific investment goals, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Investment based on this is at one's own risk

### Related Stocks

- [NDAQ.US](https://longbridge.com/en/quote/NDAQ.US.md)
- [FDS.US](https://longbridge.com/en/quote/FDS.US.md)
- [ARTY.US](https://longbridge.com/en/quote/ARTY.US.md)
- [MCO.US](https://longbridge.com/en/quote/MCO.US.md)
- [SPGI.US](https://longbridge.com/en/quote/SPGI.US.md)

## Related News & Research

- [Moody's Corporation $MCO Shares Sold by Altarock Partners LP](https://longbridge.com/en/news/298446885.md)
- [HighTower Advisors LLC Decreases Stock Holdings in Moody's Corporation $MCO](https://longbridge.com/en/news/298285665.md)
- [S&P Global Inc. $SPGI Shares Sold by Groupe la Francaise](https://longbridge.com/en/news/298284536.md)
- [BUZZ-William Blair upgrades FactSet, says AI risk overstated](https://longbridge.com/en/news/298353131.md)
- [Maestria Partners LLC Decreases Stock Holdings in Moody's Corporation $MCO](https://longbridge.com/en/news/298020071.md)

---
> **Disclaimer: This article is for reference only and does not constitute any investment advice.**