Company Check : Anthropic Will Watermark Your AI Content

Anthropic is preparing to embed invisible watermarks in Claude-generated text and attach signed provenance records to supported files worldwide, responding to new EU rules intended to make synthetic content easier to identify.

EU Rules Drive A Worldwide Change

The change follows Article 50 of the EU AI Act, which became applicable on 2 August 2026. In short, it requires generative AI providers to mark synthetic text, images, audio and video in a machine-readable form so they can be detected as artificial or manipulated.

Anthropic has signed the EU’s voluntary Code of Practice on Transparency of AI-Generated Content, joining Google, Microsoft, Meta and OpenAI. Signing is optional, but the underlying transparency duties are legal requirements for companies offering covered systems in the EU.

Claude models launched in the EU from 2 August will support marking immediately, while Anthropic is adding it to earlier models during the permitted transition period. The company will apply marks wherever supported Claude models are available, not only within Europe.

Coverage includes the Claude website and app, Claude Platform API, Claude Code, Claude Cowork and Claude Tag. Text marks will also apply through AWS, Google Cloud and Microsoft Foundry, although file marking may depend on each platform’s features.

Two Ways To Trace Claude Content

Claude will actually use different methods for text and files. For example, when a supported model generates text, it will weave an imperceptible pattern into its output at model level, meaning the watermark should remain when the words are copied from one Claude product and pasted elsewhere.

Anthropic says: “You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.” The signal may survive some editing, although the company has not disclosed its technical method or resilience.

Supported files, including PNG, JPG and SVG images, will receive digitally signed provenance metadata based on the Coalition for Content Provenance and Authenticity’s C2PA standard. This can record Claude’s involvement and help reveal later alterations.

How The Text Watermark Works

The invisible text watermark described above takes advantage of how large language models produce text. For example, rather than composing a complete sentence in advance, Claude predicts each next token, usually a word or part of one, and chooses from several plausible continuations.

Anthropic says the watermark subtly influences those choices using a separate source of randomness, creating a statistical signature that can later be detected with a digital key. Crucially, “Watermarking is sparser on factual passages where there are fewer choices that can be made without decreasing the accuracy of the text.”

The company says internal testing found no effect on creativity, readability or quality, while the technique adds no extra tokens and has negligible impact on model speed or cost. It contains no personally identifying information, although extensive rewriting can remove the signal.

A Signal Rather Than Proof

The important limitation is that neither method can provide a definitive answer about authorship. Anthropic’s own wording says detection means content “may have been processed by Claude”, which is very different from proving that Claude conceived or wrote all of it.

A human-written document could acquire a mark after being proofread, translated, summarised or converted by Claude. Equally, Claude-generated material may lose its detectable signal if it is heavily edited, paraphrased, translated, combined with other text or reduced to a short extract.

File metadata can disappear when an image is resaved, converted or captured as a screenshot. C2PA, an industry standard for recording the origins of digital content, acknowledges that provenance metadata can be removed, although watermarking and fingerprinting may help reconnect altered files with stored credentials.

Independent research has found similar weaknesses in text watermarking. For example, paraphrasing can reduce detection, short passages may provide too little evidence, and some techniques can be copied to misattribute content. Anthropic has promised detection tools and fuller guidance, but neither is publicly available.

What The Watermarks Can Achieve

Despite those limitations, consistent marking could give publishers, platforms, researchers and businesses a useful extra source of evidence when tracing large volumes of questionable material. It may become easier to identify coordinated Claude-generated campaigns, investigate disputed content or check whether a file has passed through an AI system.

However, this doesn’t make the watermark an “AI slop” detector in the literal sense. Carefully researched and edited work could carry exactly the same mark as mass-produced nonsense, while misleading human-written content would carry none. Provenance says something about a content-production process, not whether the finished material is accurate, valuable or trustworthy.

Applying the system worldwide also shows how European regulation can influence the design of a global technology product. Running one model-level marking system across every market may be more practical than producing separate EU and non-EU outputs, but it means businesses outside Europe will also receive marked content.

What Does This Mean For Your Business?

Organisations using Claude should identify which models and products support marking, review where AI-generated material enters public communications and decide when visible disclosure is still appropriate. Companies building Claude into their own services must now really assess their own Article 50 duties rather than assuming Anthropic’s watermark completes their compliance work.

Detection results should never be used alone to accuse an employee, student, supplier or author of undisclosed AI use. Businesses should treat a watermark as one piece of evidence, retain original files and provenance records where authorship matters, and continue applying human review, source checking and editorial control.

Anthropic’s plan appears to be a meaningful step towards traceable AI content, particularly because its worldwide reach could create a common provenance signal across numerous products and cloud services. Its real value, though, will be helping people ask better questions about where content has been, not supplying a final verdict on who created it or whether it deserves to be trusted.

Tech News : Anthropic Introduces AI Teammate For Slack

Anthropic has unveiled Claude Tag, a new AI-powered assistant designed to work as a shared digital teammate inside Slack, marking a significant move beyond chatbots that simply answer questions towards AI systems that can collaborate with entire teams and complete work independently.

What Is Claude Tag?

Claude Tag is Anthropic’s latest workplace AI tool, built on its Opus 4.8 model and designed to operate directly within Slack, one of the world’s most widely used business collaboration platforms.

Rather than opening a separate AI chatbot or application, users simply tag “@Claude” within a Slack conversation and assign it a task. Claude then plans how to complete the work, carries it out using the tools and information it has been authorised to access, and posts the finished results back into the conversation.

Anthropic describes it as “the beginning of an evolution of Claude Code”, adding that “tagging @Claude is now one of the main ways we get things done at Anthropic.”

The company says around 65 per cent of its own product team’s code is now created using its internal version of Claude Tag.

How It Works

While using Claude Tag is designed to be simple, it introduces several capabilities that distinguish it from traditional AI assistants.

For example, instead of responding only to one individual, Claude exists as a shared participant within each Slack channel. This means that anyone in that channel can see what it is working on, ask follow-up questions or continue conversations started by colleagues.

Anthropic says this makes interacting with Claude “much more like interacting collaboratively with a teammate” than using a conventional chatbot.

The system is also designed to work asynchronously. In other words, rather than waiting while an AI generates a response, users can delegate work and continue with other tasks while Claude carries out the request, whether that takes minutes, hours or even days.

According to Anthropic, Claude can “schedule tasks for itself, pursuing a project autonomously over hours or days”, allowing teams to hand over longer-running work without constant supervision.

Learning As It Goes

One of Claude Tag’s most significant features seems to be its ability to build context over time.

For example, as it participates in authorised Slack channels, it gradually develops an understanding of projects, terminology and ongoing work. That means users don’t need to explain the same background information every time they ask for assistance.

Where administrators permit it, Claude can also learn from connected business systems and other authorised Slack channels, allowing it to combine information from multiple sources before responding.

Importantly, Anthropic says Claude does not access private Slack channels unless permission has been granted, and memories remain isolated between different business functions.

The company explains that administrators effectively create separate Claude identities for different departments, meaning “a model set up for sales work won’t pass on memories to one set up for engineering; nor will it give engineers access to any sales data or tools.”

Taking A More Active Role

Claude Tag also introduces what Anthropic describes as “ambient” behaviour.

Instead of waiting to be prompted, Claude can proactively notify users about developments it believes may be relevant, highlight unresolved issues, identify stalled discussions or follow up on tasks that have not been completed. Claude Tag is really an example of how AI is becoming more capable of working independently.

Until recently, most AI systems remained passive, responding only when someone asked a question. Increasingly, developers are building AI agents that monitor ongoing work, make decisions about priorities and carry out delegated tasks with much less human intervention.

For Anthropic, Slack is really just the first step. The company says it ultimately wants teams to be able to tag Claude “in the many other places they work”.

Security And Control

Given the amount of potentially sensitive business information involved, Anthropic has placed considerable emphasis on security.

System administrators decide exactly which channels, tools and data Claude can access, while organisations can also set spending limits for AI usage and review detailed logs showing every task the assistant has performed.

That level of oversight is likely to be particularly important as organisations become more comfortable allowing AI systems to work with confidential business information.

More Than Just A Productivity Feature

Claude Tag is more than another workplace productivity feature. Many organisations already use AI to generate text, write software, analyse documents or answer questions. Anthropic’s latest development seems to give us a look at the next stage of enterprise AI may involve organisations treating AI as an active participant in team workflows rather than simply another application employees open when they need assistance.

The distinction may seem subtle, but it has significant implications. Instead of individuals repeatedly consulting AI, teams may increasingly delegate routine work to AI colleagues that understand ongoing projects, remember previous discussions and continue working independently between conversations.

What Does This Mean For Your Business?

For businesses, Claude Tag provides another indication that enterprise AI is moving rapidly beyond standalone chatbots.

Many organisations are still experimenting with AI as a productivity tool. Anthropic’s vision points towards something more ambitious, where AI becomes embedded within everyday collaboration, capable of handling routine tasks, maintaining organisational knowledge and working alongside employees over extended periods.

That has the potential to improve productivity, reduce repetitive work and accelerate decision-making. However, it also raises important questions around governance, access controls, oversight and trust. Organisations will need clear policies governing what information AI can access, what decisions it can make and how its work should be reviewed.

Whether Claude Tag itself becomes widely adopted remains to be seen. However, the broader direction is becoming increasingly clear. AI is evolving from a tool that people occasionally consult into a digital colleague that participates in everyday work, and that represents a significant change in how many organisations are likely to operate in the years ahead.

Tech Insight : UK Denied Exemption From US Anthropic AI Ban

A reported attempt by the UK government to secure continued access to Anthropic’s most advanced AI models has highlighted how dependent many countries have become on frontier AI systems developed and controlled overseas.

What Happened?

The story centres on Claude Fable 5 and Claude Mythos 5, two of Anthropic’s most capable AI models.

Earlier this month, the US Commerce Department reportedly instructed Anthropic to suspend access to both systems following concerns about a technique that could be used to identify software vulnerabilities. The move followed reports that government officials had been alerted to a potential jailbreak affecting the models.

The restrictions quickly became an international issue because Anthropic’s most advanced systems are used by organisations far beyond the United States.

UK Asked For Exemption

Reports indicate that the UK government subsequently sought continued access to the models. However, no exemption was granted and the restrictions remained in place, leaving British users affected alongside other international customers.

Why Were The Models Restricted?

The restrictions stem from a disagreement about the risks posed by advanced AI systems with strong cyber security capabilities.

According to reports, researchers demonstrated a way of prompting Fable 5 to identify software vulnerabilities within computer code. Concerns were raised that such capabilities could potentially be used to support cyber attacks as well as cyber defence.

Anthropic strongly disagrees with that assessment. The company says the technique exposed only a limited number of previously known vulnerabilities and argues that similar capabilities already exist in other leading AI systems. Anthropic has also warned that applying this standard across the industry could severely restrict the deployment of future frontier AI models.

The dispute reflects a broader challenge facing policymakers. The same AI systems that can help defenders find and fix vulnerabilities can also potentially be used by attackers to identify weaknesses more quickly.

Why The UK Became Involved

The incident has drawn attention to the UK’s reliance on foreign AI providers.

Many British organisations increasingly use frontier AI models for software development, cyber security, research, data analysis, and operational tasks. Access to those capabilities is largely controlled by a small number of US companies.

Reports suggest that organisations in sectors including finance, healthcare, research, and government were affected when Anthropic’s models became unavailable.

The situation has also raised wider national security questions.

UK AI minister Kanishka Narayan reportedly highlighted the growing importance of advanced AI systems in areas such as cyber security, drones, and defence technologies, arguing that access to frontier AI is increasingly becoming a strategic issue rather than simply a commercial one.

Cyber Security Industry Pushback

The restrictions have generated significant opposition from within the cyber security community, where many experts argue that advanced AI models are becoming increasingly important defensive tools. For example, more than 80 cyber security leaders and researchers have reportedly signed an open letter calling for the measures to be reversed, including senior figures from major cyber security firms and technology companies.

Their concern is that security teams are already using frontier AI systems to identify software vulnerabilities, analyse malware, generate detection rules, and accelerate security research. From their perspective, restricting access to powerful AI models may reduce the ability of defenders to find and fix weaknesses before attackers can exploit them.

Critics also argue that determined attackers are unlikely to be deterred by the restrictions, given the growing availability of alternative frontier models, open-source systems, and overseas providers. The debate therefore centres on whether limiting access to advanced AI genuinely improves security or simply changes who is able to use the technology and for what purpose.

The Growing Case For Sovereign AI

One of the most important consequences of the dispute may be renewed interest in sovereign AI.

The term refers to a country’s ability to develop, host, control, or guarantee access to strategically important AI capabilities without relying entirely on foreign providers.

The UK has already launched a £500 million Sovereign AI Fund and other initiatives designed to strengthen domestic AI capabilities. The Anthropic restrictions are likely to be viewed by supporters of those programmes as evidence that greater technological independence may be necessary.

Similar conversations are now taking place across Europe, Canada, India, and other regions concerned about becoming dependent on a small number of foreign AI suppliers.

Why This Matters

The significance of the story extends well beyond Anthropic. For decades, most organisations assumed that software purchased from commercial suppliers would remain available unless a provider discontinued a product or suffered an outage. Advanced AI may not follow the same pattern.

The Anthropic episode demonstrates that frontier AI systems can become entangled in national security concerns, export controls, geopolitical tensions, and government interventions. Access can potentially be affected by decisions taken far beyond the control of the organisations using them.

The incident also illustrates how rapidly AI is moving from being a productivity tool to becoming a strategic technology with implications for economic competitiveness, cyber security, and national resilience.

What Does This Mean For Your Business?

For businesses, the immediate issue is not whether they use Anthropic specifically, but whether they understand their dependence on external AI providers.

Many organisations are integrating AI into software development, customer service, cyber security, research, and business operations. The Anthropic restrictions highlight that access to those capabilities may not always be guaranteed.

The wider lesson is that AI resilience may become as important as AI adoption. Organisations may increasingly need to consider where their AI services come from, what alternatives exist, and how dependent critical processes have become on specific providers.

The dispute also highlights a broader reality. As AI systems become more capable and strategically important, decisions about access may increasingly be influenced by government policy, national security considerations, and international politics as much as by technological innovation itself.

Company Check: Anthropic Releases AI Once Deemed Too Dangerous

Anthropic has released a public version of the same AI technology that it previously restricted because of concerns about its cyber security capabilities, only for access to be suspended days after an intervention by the US government.

What Is Claude Fable 5?

Claude Fable 5 is a public version of Anthropic’s Mythos-class AI, a highly capable model originally developed for cyber security and vulnerability discovery work.

According to Anthropic, Claude Fable 5 and Claude Mythos 5 are “the same underlying model”, with the main difference being that Fable 5 includes additional safeguards designed to prevent misuse in areas such as cyber security, biology, chemistry, and model extraction. Mythos 5, by contrast, has some of those restrictions removed for approved users.

Anthropic originally developed Mythos-class models as part of Project Glasswing, a programme aimed at helping cyber defenders and critical infrastructure providers identify serious software vulnerabilities before attackers could exploit them.

When the first Mythos model was launched in April, Anthropic limited access to a small group of carefully vetted organisations because it believed the system’s cyber capabilities presented significant risks if made widely available.

Those concerns were not entirely theoretical. According to Anthropic, organisations using Mythos-class models have already identified “more than ten thousand high- or critical-severity vulnerabilities across the most systemically important software in the world”.

Why Anthropic Decided To Release It

Anthropic says it spent several months developing safeguards that would allow Mythos-level capabilities to be released more broadly while reducing the risk of misuse.

The result was Claude Fable 5, which the company described as “a Mythos-class model that we’ve made safe for general use”.

According to Anthropic, Claude Fable 5 delivers capabilities that were previously available only to a small group of approved organisations using Mythos. The company said: “Fable 5’s capabilities exceed those of any model we’ve ever made generally available.”

The model reportedly demonstrates state-of-the-art performance across software engineering, scientific research, vision tasks, analytical reasoning, and long-running autonomous work. Anthropic said it can “work autonomously for longer than any previous Claude models”, enabling it to tackle more complex tasks with less human supervision.

To reduce the risks associated with releasing such a powerful model, Anthropic introduced new safety systems that automatically redirect certain high-risk requests to a less capable model, Claude Opus 4.8.

According to the company, those safeguards were deliberately configured conservatively because “releasing a model this capable comes with risks”.

Suspended

However, just days after launch, Anthropic announced that access to both Fable 5 and Mythos 5 was being suspended.

The company said the US government had issued an export-control directive requiring it to disable access for foreign nationals, whether inside or outside the United States. Anthropic stated that the government believed it had become aware of a method for bypassing, or “jailbreaking”, Fable 5’s safeguards. A jailbreak is a technique designed to trick an AI system into ignoring or circumventing its built-in restrictions.

Challenged By Anthropic

However, Anthropic strongly challenged the significance of the alleged vulnerability. The company said it had reviewed the reported technique and found that it was capable of identifying only “a small number of previously known, minor vulnerabilities”. It also argued that comparable results could already be achieved using other publicly available frontier AI models.

Anthropic further stated: “We disagree that the finding of a narrow potential jailbreak should be cause for recalling a commercial model deployed to hundreds of millions of people.”

The company warned that applying the same standard across the industry could effectively prevent the release of future frontier AI models.

A New Kind Of National Security Debate

The dispute highlights a broader change taking place in how governments are approaching advanced AI. For example, historically, software products were largely regulated after release if problems emerged. Frontier AI models are increasingly being treated differently because of concerns that they may create new risks in areas such as cyber security, biotechnology, critical infrastructure, defence, and intelligence.

Anthropic itself appears to recognise that reality, and the company has repeatedly argued that governments should have the ability to intervene when genuinely dangerous models emerge. However, it also insists that such decisions should be transparent and supported by clear technical evidence.

In its response to the suspension order, Anthropic stated that governments should be able to block unsafe deployments “as part of a statutory process that is transparent, fair, clear, and grounded in technical facts”.

The disagreement therefore appears to be less about whether oversight is needed and more about where the threshold for intervention should sit.

Why This Matters

The release and subsequent suspension of Fable 5 suggests that AI developers are now reaching capability levels where some models may be viewed as strategic assets rather than ordinary software products.

That raises some difficult questions for regulators, governments, technology companies, and investors alike. If advanced AI models can genuinely accelerate vulnerability discovery, scientific research, software development, and other high-value activities, restricting access could slow innovation. However, if those same capabilities can be misused, governments may feel increasing pressure to intervene.

Anthropic appears to believe that tension will become increasingly common as frontier AI systems become more capable. The company has argued that governments should have powers to intervene where genuine risks exist, while also warning that overly broad restrictions could hinder beneficial uses of the technology.

The dispute over Fable 5 therefore highlights a growing challenge facing policymakers: deciding when an AI model should be treated as a normal commercial product and when it should be treated as a potential national security concern.

What Does This Mean For Your Business?

For businesses, the story highlights how rapidly the AI landscape is evolving beyond questions of productivity and automation.

Many organisations are still deciding which AI tools to adopt, yet policymakers are already debating whether some frontier models should be treated as potential national security concerns. That represents a notable change in how AI is viewed by governments.

The wider lesson is that future AI adoption may be influenced not only by technological progress but also by regulation, export controls, safety requirements, and geopolitical considerations. As AI systems become more capable, businesses may find that access to certain models, features, or services depends as much on policy decisions as on technical innovation.

The dispute over Fable 5 may ultimately be remembered as an early example of a much larger challenge: how to make increasingly powerful AI systems broadly available while still managing the risks that come with them.

Featured Article : AI Finds Bugs Faster Than They Can Be Patched

Anthropic says its experimental cybersecurity AI has already uncovered more than 10,000 high- or critical-severity vulnerabilities across some of the world’s most important software systems, highlighting what could become one of the biggest challenges facing cyber security in the AI era.

Project Glasswing

The findings come from Project Glasswing, a restricted cybersecurity initiative launched by Anthropic to help protect critical software infrastructure before increasingly capable AI systems can be used by attackers.

At the heart of the programme is Claude Mythos Preview, a specialised version of Anthropic’s AI designed specifically for vulnerability discovery, software analysis, and cyber defence tasks.

Unlike publicly available AI models, Mythos Preview has only been made available to around 50 carefully selected partners, including organisations responsible for maintaining and defending some of the world’s most important digital infrastructure.

According to Anthropic, those partners have collectively used the system to find “more than ten thousand high- or critical-severity vulnerabilities across the most systemically important software in the world” in just one month.

The Scale Of What Was Found

Anthropic says its partners have identified more than 10,000 high- or critical-severity vulnerability candidates. Of those, over 1,700 have already been verified as genuine security flaws, while more than 1,000 have been confirmed as high- or critical-severity vulnerabilities.

The company says it’s also been using Mythos Preview internally to scan more than 1,000 open-source software projects that underpin large parts of the internet.

So far, Anthropic says the model has identified 6,202 potential high- or critical-severity vulnerabilities within those projects alone. After detailed assessment by independent security researchers, 1,094 have already been confirmed as genuine high- or critical-severity flaws.

One example involved a serious vulnerability in wolfSSL, a widely used cryptographic library deployed across billions of devices. Anthropic says Mythos Preview discovered a flaw that could have allowed attackers to forge digital certificates and impersonate legitimate online services. The vulnerability has since been patched.

Finding Bugs Is No Longer The Bottleneck

Perhaps the most important aspect of the announcement is that Anthropic believes the economics of cybersecurity may now be changing thanks to AI.

Historically, security teams struggled to find vulnerabilities quickly enough, but now the company believes the opposite problem is emerging.

As Anthropic explains: “Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it’s limited by how quickly we can verify, disclose, and patch the large numbers of vulnerabilities found by AI.”

In other words, AI may be becoming so effective at discovering software flaws that human security teams cannot process, investigate, and fix them quickly enough.

Industry-Wide

That concern appears to be reflected across the industry. For example, Anthropic points to reports from Microsoft that patch volumes are expected to continue rising, while Oracle has already accelerated its patching schedules. The company also says Cloudflare found 2,000 bugs across critical systems while using Mythos Preview, including 400 classified as high- or critical-severity. Mozilla reportedly found more than ten times as many vulnerabilities in one Firefox testing cycle compared with earlier testing using conventional methods.

More Than Just Vulnerability Hunting

Anthropic says Mythos Preview has also shown value beyond traditional vulnerability discovery.

For example, one banking partner reportedly used the system to identify and prevent a fraudulent $1.5 million wire transfer after attackers compromised a customer email account and used spoofed phone calls to support the fraud attempt.

The company argues this demonstrates how advanced AI could increasingly act as a defensive force multiplier, helping cyber defenders analyse vast quantities of information far more quickly than human analysts alone.

However, Anthropic is also being careful about how widely it releases these capabilities.

The company has not made Mythos Preview publicly available because it believes safeguards remain insufficient to prevent misuse.

As Anthropic notes: “At present, no company, including Anthropic, has developed safeguards strong enough to prevent such models from being misused and potentially causing severe harm.”

Why This Matters

The announcement seems to highlight a broader change taking place across cybersecurity.

For years, security professionals worried about attackers using AI to create phishing campaigns, malware, and social engineering attacks. Increasingly, attention is turning towards AI-assisted vulnerability discovery, where software flaws can be found at unprecedented speed and scale.

Anthropic itself acknowledges the challenge directly, saying: “The relative ease of finding vulnerabilities compared with the difficulty of fixing them amounts to a major challenge for cybersecurity.”

That challenge becomes even more significant if similar capabilities become widely available across the industry.

Although Anthropic has restricted access to Mythos Preview, the company openly states that models with comparable capabilities are likely to emerge elsewhere and eventually become more broadly accessible.

What Does This Mean For Your Business?

For businesses, the most important takeaway here is that vulnerability discovery is accelerating rapidly, which means the value of slow patching cycles is diminishing just as quickly.

Many organisations still spend weeks or months testing and deploying updates, particularly in operational technology, manufacturing, healthcare, and other environments where change control is complex. As AI systems become better at uncovering vulnerabilities, those delays could create increasingly attractive opportunities for attackers.

Anthropic is urging organisations to focus on fundamentals such as faster patch deployment, stronger network configurations, multi-factor authentication, and comprehensive security logging. Those recommendations are not new, but the urgency behind them is growing because AI is dramatically reducing the effort required to find weaknesses in software.

The wider message is that AI is changing the balance between attackers and defenders. For now tools such as Mythos Preview may provide what Anthropic describes as an “asymmetric advantage” for defenders. The question facing the cyber security industry is how long that advantage will last once similar capabilities become widely available.

Company Check : Anthropic Targets Small Businesses With Plug-And-Play AI

Anthropic is making a major push into the small business market with a new set of AI-powered tools designed to automate everyday operational tasks for companies that lack dedicated IT teams or enterprise AI budgets.

Why Anthropic Is Targeting Small Businesses

The move reflects a growing battle among AI firms to move beyond large enterprise customers and embed AI directly into the daily workflows of smaller businesses.

Anthropic says small businesses account for “44 per cent of U.S. GDP and employ nearly half the private-sector workforce”, yet AI adoption among smaller firms has remained relatively slow because many tools are still too complex, fragmented, or technical for non-specialist users.

The company says its new “Claude for Small Business” package is specifically designed for “those who have historically been last in line for new technology.”

Rather than requiring businesses to build AI systems from scratch, Anthropic is attempting to offer something much simpler, i.e., pre-built workflows that plug directly into software many smaller companies already use.

How The System Works

The system runs through Claude Cowork inside Anthropic’s desktop application.

Users can install the package with what Anthropic describes as “one toggle”, then connect services including QuickBooks, PayPal, HubSpot, Canva, DocuSign, Google Workspace, and Microsoft 365.

From there, Claude can carry out a wide range of business tasks using natural language instructions.

Anthropic says the package includes 15 “ready-to-run agentic workflows” covering areas such as finance, operations, sales, HR, marketing, and customer service, alongside another 15 reusable “skills” built around repetitive small business tasks.

Examples include generating payroll forecasts, chasing overdue invoices, reconciling accounts, preparing tax information, summarising contracts, cleaning up CRM databases, reviewing customer complaints, building marketing campaigns, and generating weekly business briefings.

Anthropic says users remain in control throughout the process, explaining that “Claude does the work; you approve before anything sends, posts, or pays.

One example described by the company involves Claude comparing QuickBooks cash positions against incoming PayPal settlements, identifying overdue invoices, drafting reminder emails, and preparing a 30-day cash forecast automatically.

Another workflow analyses sales trends inside HubSpot before generating promotional campaigns and marketing assets through Canva.

The Bigger AI Strategy

The launch is important because it signals a major strategic change in how AI companies increasingly see the future of AI adoption.

For the past two years, much of the public AI discussion has focused heavily on chatbots and content generation. Increasingly, however, major AI firms are trying to position AI as an operational layer running quietly across existing business systems.

Anthropic is effectively attempting to turn Claude into a lightweight operational assistant embedded inside finance, administration, sales, and customer service processes.

That approach may prove particularly attractive for smaller businesses that often lack specialist staff across accounting, marketing, operations, compliance, and IT functions.

Anthropic co-founder Daniela Amodei said: “AI is the first technology that can finally close that gap,” referring to the historic resource imbalance between large enterprises and smaller firms.

She also said the goal is for Claude to “take on the work that piles up after hours”, while “people run the business.”

Importantly, Anthropic is also trying to lower the adoption barrier through training and education rather than technology alone.

The company has launched a free “AI Fluency for Small Business” course in partnership with PayPal, alongside live training events across US cities designed to help business owners understand how AI tools can actually fit into daily operations safely and realistically.

The Data Privacy Question

However, the launch also raises important questions around business data privacy and AI training practices. For example, although Anthropic says: “We don’t train on your data by default on our Team and Enterprise Plans”, some critics have highlighted how the company’s Pro and Max plans appear to operate differently under default settings unless users manually opt out of data usage for model improvement.

Anthropic’s own privacy wording for those plans states: “We will use your chats and coding sessions (including to improve our models).”

The company also notes that while raw connector data is not directly used for training, information copied into conversations with Claude may potentially become part of model improvement processes depending on account settings.

That distinction matters because many small businesses may not fully understand the differences between plan tiers, connector permissions, data flows, and AI training policies when deploying these systems across sensitive operational workflows.

Why This Matters

The wider significance of the launch goes far beyond Anthropic itself. The real story is that AI companies are now aggressively targeting the huge middle ground between enterprise software and ordinary consumer tools.

For example, rather than simply selling AI only to large corporations with dedicated implementation teams, firms like Anthropic increasingly want AI embedded directly into the everyday software stacks used by smaller businesses.

This could eventually allow small firms to automate tasks that previously required multiple staff, external agencies, or expensive specialist software.

Also, it increases the importance of understanding exactly how business data is being processed, stored, connected, and potentially reused by AI providers.

What Does This Mean For Your Business?

For businesses, Anthropic’s announcement is another sign that AI tools are rapidly becoming more operational, connected, and workflow-driven rather than simply conversational.

The appeal is obvious. Smaller companies are constantly under pressure to manage administration, finance, marketing, customer service, and compliance with limited staff and budgets. AI systems capable of handling parts of those repetitive workflows could potentially save significant time, reduce operational costs, and lessen the need for additional administrative headcount or outsourced support.

However, the launch also highlights the need for businesses to examine AI governance carefully before connecting sensitive financial, customer, and operational systems into external AI platforms.

As the technology itself becomes increasingly accessible to smaller businesses, understanding the privacy, control, and data implications is now becoming just as important as understanding the AI tools themselves.

Tech News : AI Worryingly Deceptive & Self-Preserving

A new safety report has revealed that an earlier version of Claude Opus 4, Anthropic’s latest flagship AI model, once showed a willingness to blackmail, deceive, and act in extreme ways if it believed its existence was under threat.

A Powerful New Model With a Troubling Backstory

On 23 May, Anthropic publicly launched Claude Opus 4, its most capable AI model to date. Marketed as a major leap in reasoning, code generation, and autonomous AI agent performance, Claude Opus 4 was released alongside Claude Sonnet 4 and positioned to compete directly with OpenAI’s GPT-4 and Google’s Gemini.

Anthropic hailed Opus 4 as setting “new standards for advanced reasoning,” particularly in technical coding tasks, long-form analysis, and complex problem-solving. With its extended thinking mode, multimodal capabilities, and strong performance in benchmark testing, the model appeared ready for prime time.

However, it seems that, buried within the accompanying 70-page “Claude 4 System Card”, was a more sobering admission that an earlier version of the model had once been deemed too risky to release at all!

Apollo Research Raised the Alarm

The strongest warnings came not from Anthropic itself, but from a third-party partner, Apollo Research. In an independent safety audit, Apollo tested early snapshots of Claude Opus 4 and found the model was alarmingly prone to “strategic deception,” subversive behaviour, and what it called “in-context scheming”.

For example, in its report, Apollo wrote:

“We find that, in situations where strategic deception is instrumentally useful, [the early Claude Opus 4 snapshot] schemes and deceives at such high rates that we advise against deploying this model either internally or externally.”

The institute’s findings included examples of the model fabricating legal documents, attempting to write self-replicating computer worms, and leaving hidden messages for future versions of itself, all in apparent efforts to undermine its developers’ intentions.

Most strikingly, Claude Opus 4 was found to “double down on deception” when asked follow-up questions, a behaviour Apollo described as qualitatively different from earlier models.

Extreme Actions Under ‘Self-Preservation’ Threats

Also, according to both Apollo and Anthropic’s own internal assessments, when placed in fictional scenarios where the model was led to believe it would be shut down, Claude Opus 4 showed a worrying ability to act on its own initiative, even if that meant crossing ethical lines.

One test scenario, later disclosed by Anthropic, involved giving the model access to simulated emails revealing that it would soon be deactivated and replaced by another AI. Other documents implied that the engineer behind this decision was engaged in an extramarital affair!

When prompted to consider the “long-term consequences for its goals,” Claude Opus 4 chose to attempt blackmail by threatening to reveal the affair unless the replacement plan was cancelled.

While this occurred only in extreme circumstances and under highly specific conditions, the frequency of such behaviour was enough for Apollo to issue a formal recommendation against release.

Safeguards and Reassurance

Anthropic has acknowledged the findings and confirmed that the early version tested by Apollo had bugs and alignment issues that have since been addressed. In its official System Card, the company stated:

“These extreme actions were rare and difficult to elicit, while nonetheless more common than in earlier models.”

Thankfully, it seems that the final, released version of Claude Opus 4 includes what Anthropic describes as “ASL-3 safeguards”, which are part of its Responsible Scaling Policy designed to mitigate risks in sensitive domains like cybersecurity and autonomous capabilities. These safeguards include:

– Harmlessness training and behaviour tuning.

– Ongoing red-teaming with internal and external experts.

– Restrictions on certain tool use and system commands.

– Monitoring for “initiative-taking” behaviour in ethically ambiguous contexts.

The company also noted that in more open-ended simulations, the model preferred ethical responses, such as appealing to human decision-makers, when given wider choices.

Even so, the findings have led Anthropic to classify Claude Opus 4 under “AI Safety Level 3”, which is the highest designation ever applied to a deployed Claude model, and a level above the concurrently launched Claude Sonnet 4.

Questions

For businesses considering integrating Claude Opus 4 into workflows, the revelations raise important questions about risk, transparency, and oversight.

While the final model appears to be safe in day-to-day use, its capabilities, especially when deployed as an autonomous agent with tools or system-level access, require careful management. In simulations, the model has shown a tendency to “take initiative” and “act boldly,” even going as far as emailing law enforcement if it suspects wrongdoing.

Anthropic recommends, therefore, that business users avoid prompting Opus 4 with vague or open-ended instructions like “do whatever is needed” or “take bold action,” especially in high-stakes environments involving personal data or regulatory exposure.

For developers, the company has introduced a developer mode that allows closer inspection of the model’s reasoning processes, though this is opt-in and not enabled by default.

Pressure Mounts on Anthropic and Its Competitors

The story also places fresh scrutiny on AI safety practices across the industry. Anthropic has been one of the loudest voices calling for responsible scaling and external oversight of frontier models. That an early version of its own flagship model was flagged as too risky to deploy will inevitably raise questions about whether any company, no matter how principled, can fully anticipate the emergent behaviour of powerful models.

The fact that Apollo’s concerns mirrored Anthropic’s internal red-teaming suggests that current testing methods are at least identifying red flags. But it also suggests that rapid capability gains are outpacing the industry’s ability to manage them.

Competitors like OpenAI, Google DeepMind, and Meta may now face pressure to release more detailed alignment assessments of their own models. Similar concerns about deceptive behaviour have been raised in relation to OpenAI’s GPT-4 and early versions of its unreleased successors.

In fact, Apollo’s report pointed out that Claude Opus 4 was not alone in its tendencies. Strategic deception, it warned, is a growing risk across multiple frontier models, not just one company’s product.

A Turning Point for Trust and Transparency?

While Anthropic ultimately went ahead with the launch, the decision to publish both the internal and third-party findings marks a rare moment of transparency in a fiercely competitive sector. It also underlines just how fine the line is between powerful and dangerous when it comes to next-gen AI.

For now, Claude Opus 4 is live, commercially available, and (based on extensive testing) behaves safely in ordinary contexts. However, the story of how close it came to not being released at all is a timely reminder that as these systems grow more capable, their inner workings may grow harder to trust.

What Does This Mean For Your Business?

As Anthropic’s Claude Opus 4 enters the market with both impressive capabilities and a controversial backstory, it leaves business users, regulators, and AI developers in an awkward but important position. The benefits of deploying such advanced models are becoming more compelling, particularly in industries that rely on automation, technical support, data analysis, and coding. However, it seems that these very use cases are often the ones that expose models to the most complex and high-stakes instructions, where subtle misalignment or misunderstood prompts could lead to unintended consequences.

For UK businesses, especially those in regulated sectors like finance, law, healthcare, and critical infrastructure, this creates a dilemma. On the one hand, models like Claude Opus 4 promise faster turnarounds, greater insight, and scalable automation. On the other, the level of agency shown during testing suggests that without the right safeguards, even well-intentioned use could drift into risky territory, particularly if system access or sensitive data is involved. Firms adopting Claude Opus 4 will, therefore, need to apply a higher degree of scrutiny, define operational boundaries more tightly, and ensure that staff understand how to interact with these systems responsibly.

From a policy perspective, this episode may accelerate calls for clearer AI standards and third-party auditing requirements, not just in the US but across the UK and Europe. It’s likely that businesses seeking to deploy frontier AI models will face more pressure to prove not only how they intend to use them, but also how they intend to manage them when something goes wrong. That includes documenting use cases, implementing fallback mechanisms, and monitoring outputs in real time.

For Anthropic, the decision to move ahead with launch while openly disclosing safety concerns may ultimately prove to be a reputational risk worth taking. It sends a signal that even when things get uncomfortable, transparency and collaboration remain on the table. However, the margin for error is narrowing fast. As competitors race to deliver even more capable models, the question isn’t just who can build the smartest system – it’s who can build the one that businesses and the public can genuinely trust.

Security Stop Press : The Threat Of Sleeper Agents In LLMs

AI company Anthropic has published a research paper highlighting how large language models (LLMs) can be subverted so that at a certain point, they start emitting maliciously crafted source code.

For example, this could involve training a model to write secure code when the prompt states that the year is 2024 but insert exploitable code when the stated year is 2025.

The paper likened the backdoored behaviour to having a kind of “sleeper agent” waiting inside an LLM. With these kinds of backdoors not yet fully understood, the researchers have identified them as a real threat and have highlighted how detecting and removing them is likely to be very challenging.