Skip to main content

Claude Opus 4.8: Anthropic makes a more honest AI

An Anthropic logo behind an iPhone logging into Claude

Anthropic isn't ready to let regular users look at its supposedly super-powerful Claude Mythos AI model just yet. But the AI company has just released an upgrade to its flagship product, Claude Opus — now in its 4.8 version.

"It builds on Opus 4.7 with improvements across benchmarks, and is a more effective collaborator," Anthropic promised in a press release Thursday. Indeed, the benchmark numbers, below, show very minor improvements across the board.

One major improvement, allegedly, is in the area of hallucinations. Claude Opus 4.8 won't lie to users as much. "Early testers report that Opus 4.8 is more likely to flag uncertainties about its work and less likely to make unsupported claims," Anthropic said, touting the model's "honesty."

Claude Opus 4.8 has 'better judgment'

"Claude Opus 4.8 has noticeably better judgment," an engineer at Shopify, Tom Pritchard, told Anthropic. The coding version of the model "asks the right questions, catches its own mistakes, and pushes back when a plan isn't sound."

Given the increasing number of horror stories about AI agents deleting entire corporate databases, that promise may be music to the ears of vibe coders everywhere.

To please power users, Anthropic is offering a significant discount on "fast mode," where Claude will work at 2.5 times regular speed. Fast mode "is now three times cheaper than it was for previous models," the company said.

Users on Reddit weren't buying it, however. Many feared a loss of access to a more popular model, Claude Opus 4.6. "Nobody trusts the benchmark charts," wrote one redditor in summary, noting that Opus 4.7 also seemed to have some pretty good numbers when it was released.

Whether or not we can trust the benchmarks — and to be clear, Mashable hasn't independently verified these numbers — here's what Anthropic is claiming.

A list of benchmark numbers for Claude Opus 4.8
Credit: Anthropic

How to try Claude Opus 4.8

Claude Opus 4.8 is available now via Anthropic's website, Claude.AI, as well as via the Claude API, plus Anthropic partners like Microsoft Foundry.

The new model is priced exactly the same as its predecessors, which is to say models going all the way back to Claude Opus 4.5. All of them will cost you $5 per million input tokens and $25 per million output tokens.

Given that Anthropic is promising Claude Mythos within a matter of weeks, however, you may want to hang back and wait to see whether that model can be even more "honest" about its hallucinations.



from Mashable https://ift.tt/Hrsn34b
via IFTTT

Comments

Popular posts from this blog

The Nintendo Switch has been the US’s bestselling console for 23 straight months

Photo by James Bareham / The Verge It’s been a good two years for the Nintendo Switch. According to Nintendo, the gaming tablet has been the bestselling console in the US for 23 straight months. And according to data from the NPD Group, it just had its best October ever, moving 735,926 units of both the Switch and Switch Lite in the US. The company says that represents a 136 percent increase compared to last year. To date, the Switch has sold 22.5 million units in the US, and last week Nintendo revealed that more than 68 million units have been sold globally . “We’re excited about our momentum,” says Nick Chavez, Nintendo of America’s SVP of sales and marketing. Chavez puts the company’s big October down to two main factors. One is a better supply of stock; this year in particular, it’s often been hard to find a Switch on store shelves. This has only been exacerbated by increased demand due to a combination of the pandemic and the breakout success of Animal Crossing: New Horizons . ...

Instagram accidentally reinstated Pornhub’s banned account

After years of on-and-off temporary suspensions, Instagram permanently banned Pornhub’s account in September. Then, for a short period of time this weekend, the account was reinstated. By Tuesday, it was permanently banned again. “This was done in error,” an Instagram spokesperson told TechCrunch. “As we’ve said previously, we permanently disabled this Instagram account for repeatedly violating our policies.” Instagram’s content guidelines prohibit  nudity and sexual solicitation . A Pornhub spokesperson told TechCrunch, though, that they believe the adult streaming platform’s account did not violate any guidelines. Instagram has not commented on the exact reasoning for the ban, or which policies the account violated. It’s worrying from a moderation perspective if a permanently banned Instagram account can accidentally get switched back on. Pornhub told TechCrunch that its account even received a notice from Instagram, stating that its ban had been a mistake (that message itse...

MVP versus EVP: Is it time to introduce ethics into the agile startup model?

Anand Rao Contributor Share on Twitter Anand Rao is global head of AI at PwC . The rocket ship trajectory of a startup is well known: Get an idea, build a team and slap together a minimum viable product (MVP) that you can get in front of users. However, today’s startups need to reconsider the MVP model as artificial intelligence (AI) and machine learning (ML) become ubiquitous in tech products and the market grows increasingly conscious of the ethical implications of AI augmenting or replacing humans in the decision-making process. An MVP allows you to collect critical feedback from your target market that then informs the minimum development required to launch a product — creating a powerful feedback loop that drives today’s customer-led business. This lean, agile model has been extremely successful over the past two decades — launching thousands of successful startups, some of which have grown into billion-dollar companies. However, building high-performing product...