Agentic OS: 16 AI Assistants Ranked by Tier (2026)

18 min readBy Nathan House

In the space of a few weeks, the big AI labs all started selling the same thing. xAI launched Grok Bot on 11 August. Meta launched Muse on 8 September. Manus launched Cue on 28 September, and OpenAI announced dots at its developer day the next day. Google's Gemini Spark and Anthropic's Claude Cowork were already out earlier in the year. Each one is an assistant that remembers you, runs around the clock on its own cloud computer and does real tasks across your apps.

That idea has a name: an agentic OS, an AI assistant you build up over time. Until now, you mostly had to build one yourself, from free open-source projects with hundreds of thousands of GitHub stars to paid communities and Anthropic's own Claude Projects. Now you can also just sign up for one. Most guides list all of these as if they were the same thing. They aren't.

In this guide, we rank 16 of them using the four-tier model we teach in our AI Master's program, and we show the evidence for every placement. They're listed from the lowest tier to the highest, and each one is marked as either something you build or run yourself (10 of them) or a ready-made assistant from a big lab that needs no technical setup (6). Then we look at where it's all heading: from one person's assistant to a whole company's shared brain.

⏱️ TL;DR — if you've only got 30 seconds

An agentic OS is an AI assistant with memory, an identity and skills that compound over time. This year Google, xAI, Meta and OpenAI all started selling one ready-made, but most aren't available in the UK yet. We rank 16 of them on four tiers: harness (the base tool), personal AI (it knows you), AI infrastructure (it runs your operation), and shared AI infrastructure (a whole team's brain). Almost everything today is single-user — even Anthropic's new Projects ("a project belongs to one user"). The real shift, and what Y Combinator is now funding, is AI going multiplayer: one shared company brain instead of a thousand private threads.

What Is an Agentic OS?

Think about the difference between a temp and a long-serving assistant. Both are smart. The temp needs everything explained every morning. The assistant already knows how you like your reports, which clients are difficult and what went wrong last time, and just gets on with it.

A plain AI chat is the temp. Close the window and it forgets you. An AI coding tool like Claude Code is further along, it can even keep notes between sessions, but on its own it still doesn't know your standards, your projects or how you work. An agentic OS is what you get when you give it the long-serving assistant's advantages: a memory that lasts between sessions, a profile of who you are and how you work, and skills, which are written-down procedures it can reuse, so it does a task your way every time.

Nobody owns the name. You'll see "agentic OS", "personal AI", "personal AI infrastructure", "AI assistant" and "second brain" used for the same idea. Slack uses "agentic OS" for something bigger again: a layer that coordinates many agents across a whole company.

That spread of meanings is exactly why a single list doesn't work. A tool that remembers your writing style and a platform that runs a business's agents both get called an agentic OS, so we need a way to tell them apart.

The Four Tiers of AI Assistants

We use four tiers. Each one is defined by what the system does for you, not by how many features it lists.

The four tiers as an iceberg: Tier 1 Harness at the visible tip, then Tier 2 Personal AI, Tier 3 AI Infrastructure and Tier 4 Shared AI Infrastructure widening below the water

Tier 1 is the harness. Claude Code, Codex, Cursor, Gemini CLI. A language model with tools, so it can read files, run commands and browse. That's what turns a model into an agent. It's powerful, and it's where most people work, but on its own it doesn't know you: your standards, your projects, the decisions you've already made.

Tier 2 is personal AI. You add a memory, an identity and skills to the harness. Now it's "a tool that knows you", and it gets a little better every time you use it.

Tier 3 is AI infrastructure. The assistant stops being something you chat to and becomes "a system that runs your operation." It connects to the real systems a business depends on, such as servers, billing, publishing and monitoring, and it has a deliberate security model so it can do that without causing damage.

Tier 4 is shared AI infrastructure. The same system, shared by a team. Everyone works from one shared base of context, skills and standards, and each person has their own space with their own permissions.

To check where each product belongs, we scored all 16 against the seven capabilities we teach: identity, memory, skills, orchestration (coordinating several agents), execution (working while you're away), interface (how you reach it) and governance (the rules that stop it doing damage). We read each project's own documentation and code, and every placement below rests on a quote from its own docs.

One thing surprised us. Nearly every Tier 2 product already has the Tier 3 features: sub-agents, scheduled jobs and permission rules. Most of those come free with the harness underneath. So features alone don't decide the tier. What does is reach: what the system is actually trusted to run, and for how many people.

The 16 agentic OS products placed on the four tiers: gstack on Tier 1; LifeOS, NanoClaw, Hermes, Simon Agentic OS and OpenClaw, plus the ready-made Gemini Spark, Meta Muse, Cue, OpenAI dots, Claude Cowork and Grok Bot, on Tier 2; Claude Projects, HAL and Paperclip on Tier 3; Hosted HAL on Tier 4

16 Agentic OS Options, Ranked by Tier

Star counts are from GitHub on 2 October 2026. A star is a developer bookmarking a project, so it is a rough measure of popularity, not quality. "MIT licence" means the code is free to use and change.

Capability key: ✓ yes · ◐ partial · ✗ no — scored from each product's own docs. For the six ready-made assistants, every tick rests on a quote from the vendor's own pages; anything the vendor hasn't documented is marked ✗.

Each card is marked with a chip showing which kind it is. Build it yourself means you install, run or assemble it: you need a terminal, an API key, or at least a Claude Code subscription and the patience to set it up. Ready-made means one of the big labs runs it for you: you sign up in an app, connect your email and calendar, and it starts working.

The ready-made ones are new: Google, xAI, Meta, Manus, OpenAI and Anthropic all sell one now, most of them launched since August. Every one lands on Tier 2. Grok Bot is the only one already reaching toward Tier 4, because a whole team can share a bot. None of them is trusted to run an operation yet, which is what Tier 3 takes.

1. gstack

Build it yourself Tier 1 add-on
A real gstack session in Claude Code, from Y Combinator's video: gstack's office-hours skill asking a founder for the strongest evidence that anyone wants their product

Real screen. Source: Y Combinator video with Garry Tan (4:52).

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

gstack is Garry Tan's personal Claude Code setup, published as open source (MIT licence, 134,729 stars). Tan is the president and CEO of Y Combinator, the startup accelerator. The pack describes itself as "Twenty-three specialists and eight power tools, all slash commands": a CEO reviewer, a designer, a QA tester, a security officer and so on, each a role Claude Code can play on demand.

It's a brilliant way to see how a heavy user structures their skills. But it's a skill pack for one developer, not an assistant that knows you. It sits on top of Tier 1 and makes it sharper.

Official project: github.com/garrytan/gstack

Best for: developers who want a proven set of review and planning skills to copy.

2. LifeOS, formerly PAI

Build it yourself Tier 2
LifeOS's real Pulse dashboard, from the LifeOS site: a live board of agent runs, a progress chart and a list of claims

Real screen. Source: LifeOS site.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

Daniel Miessler's Personal AI Infrastructure was renamed LifeOS in July 2026 (v6.0.0, MIT, 19,265 stars). It's the closest in spirit to how we build: it's designed as "a single Digital Assistant" built around you, with your identity and goals loaded at the start of every session, a memory of past decisions, and a denylist that "blocks dangerous operations regardless of what any prompt says."

Miessler comes from the security world, and it shows. LifeOS treats safety as part of the design rather than an afterthought.

Official project: github.com/danielmiessler/LifeOS

Best for: people who want a deeply personal assistant for their whole life, not just work, and are happy running it on Claude Code.

Official video from the maker.

3. NanoClaw

Build it yourself Tier 2
NanoClaw's real terminal installer, from the NanoClaw blog: the bash nanoclaw.sh setup running with the NanoClaw ASCII logo

Real screen. Source: NanoClaw blog.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

NanoClaw (MIT, 30,867 stars) is a smaller, security-first take on the same idea. Every agent runs inside its own container, a sealed-off box on the computer that it cannot see out of, so "Bash access is safe because commands run inside the container." It connects to WhatsApp, Telegram, Slack, Teams and iMessage, keeps a file-based memory, and runs scheduled tasks.

It describes itself as "Built for the individual user." It does have user roles, so a few people can share an agent, but its own docs warn that between people in the same group, "Information will cross-pollinate through agent memory."

Official project: github.com/nanocoai/nanoclaw

Best for: anyone who wants a chat-app assistant and cares about containment.

Official video from the maker.

4. Hermes Agent

Build it yourself Tier 2
Hermes Agent's real web dashboard, from its official docs: a Kanban board of agent tasks in triage, to do, ready, in progress, blocked and done

Real screen. Source: Hermes Agent docs on GitHub.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

Hermes Agent from Nous Research (MIT, 250,661 stars) is the one with the most interesting memory. It keeps "bounded, curated memory that persists across sessions" plus a user profile with "your preferences, communication style, expectations", and it writes new skills for itself after complex tasks. It runs from a laptop or a $5 server and talks to you through Telegram, Discord, Slack, WhatsApp and Signal.

Its security policy is refreshingly direct: "Hermes Agent is a single-tenant personal agent." Several people can talk to the same instance, but it doesn't keep them apart.

Official project: github.com/NousResearch/hermes-agent

Best for: people who want an assistant that learns and improves on its own, and who want to see a memory system worth copying.

5. Simon Scrapes' Agentic OS

Build it yourself Tier 2
Illustration of Simon Scrapes' Agentic OS: its Command Centre dashboard

Illustration.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

The Agentic OS from Simon Scrapes is the paid option: a private repository you get through his Agentic Academy community (a monthly or annual membership). It's built on Claude Code for business owners rather than developers, with brand files, a personality file, a curated memory, 26 skills covering marketing, strategy and operations, scheduled jobs that "run autonomously" since version 1.1.4, and a local Command Centre dashboard.

It handles several clients from one install, but that's one operator managing many clients, not many people sharing it. His "Team OS" is in beta, and we haven't been able to verify how it works, so its multi-user and permissions marks are unconfirmed.

Best for: non-developers who want a finished, supported system rather than building their own.

Official video from the maker.

6. Gemini Spark

Ready-made Tier 2
Gemini Spark's real desktop app, from Google's Spark page: Spark replying that it has set up a weekly schedule to track interior design internships every Monday morning

Real screen. Source: Google's Gemini Spark page.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

Google announced Gemini Spark at its I/O conference on 19 May. It started with testers and a beta for Google AI Ultra subscribers in the US, and now comes with Google AI Pro ($19.99 a month) and Ultra. Google calls it "Your 24/7 personal AI agent", and it works in the background "even if your phone and laptop are turned off." It connects to Gmail, Calendar, Drive, Docs and other Google apps (those connections are off until you turn them on), and you can save reusable instructions for jobs you repeat.

It's designed to check with you before major actions. But Google's own help page is candid about the limit: if a scheduled task runs while you're offline, you may not be able to stop an unintended action. It isn't offered in the UK, the European Economic Area or Switzerland.

Official page: gemini.google/overview/agent/spark

Best for: people who live in Gmail, Calendar and Docs and want an assistant with no setup.

Official video from the maker.

7. Meta Muse

Ready-made Tier 2
Meta Muse's real phone app, from Meta's launch post: a checkout card where Muse asks permission to place an $80 order, with Deny and Allow buttons and its browser blocked while it waits

Real screen. Source: Meta's Muse launch post.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

Meta launched Muse in the US on 8 September, on iOS, Android, muse.ai and WhatsApp, and it's now in the US and Canada. It's free with a usage limit, or $20 a month (Power) and $100 a month (Maximum) for more. It's the most clearly aimed at non-technical people of the six.

It also has the most thought-through safety design. Each person's Muse runs on a "Muse Secure VM", a dedicated virtual machine that holds the agent and the person's data. And "A separate Sentinel agent runs on that same machine, kept apart from Muse at the system level": nothing Muse does reaches the internet unless the Sentinel approves it. Under the hood it does more than it shows: Meta's security page says it "launches swarms of subagents" and builds its own tools. A small-business version, launched on 29 September, connects Muse to a company's Facebook Pages, Meta ad accounts, Shopify and more.

Official announcement: about.fb.com

Best for: non-technical people in the US or Canada who want the strongest guardrails of the ready-made options.

Official video from the maker.

8. Cue

Ready-made Tier 2 · provisional
Manus Cue's real app, from cue.im: the desktop agent list and a Senior Designer agent's chat beside the same chat on a phone

Real screen. Source: cue.im.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

Manus launched Cue on 28 September, a new app for your personal agents on phone and desktop. It goes further than the others in giving an agent a life of its own: "In Cue, each agent has its own email, phone number, wallet, and computer, so it can send messages, pay within the budget you set, and see a task through on its own machine." You choose what it can access and which actions need your approval.

You can also "Put several agents in a group chat with a shared goal, and they hand work to each other." It's in early access and free with an invite code.

Why provisional: Cue has no help or documentation pages yet. Manus hasn't published how its memory works, which model it uses or where it's available, so memory is marked ✗ (not documented) and we'll revisit this placement once it has.

Official announcement: manus.im/blog/introducing-manus-2-0

Best for: early adopters who want an agent that can message and pay on their behalf, inside a budget.

Official video from the maker.

9. OpenAI dots

Ready-made Tier 2
A real OpenAI dot, from OpenAI's announcement: a dot chat updating a study, beside the Gmail draft with charts it wrote

Real screen. Source: OpenAI's dots announcement.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

OpenAI announced dots at its developer day on 29 September. In its words: "Powered by GPT‑6 Astra, they have their own cloud computer, learn from feedback over time, and can work towards your goals 24/7." Your first dot "is included in your Pro or Business Premium plan at no extra cost", and ChatGPT Pro starts at $100 a month.

A dot builds on your existing ChatGPT memories, connects to your apps through plugins, and works in ChatGPT, Slack and Teams. Custom Rules decide how far it goes on its own, from acting without asking to handing the task back to you. It's rolling out "in markets excluding the European Economic Area, Switzerland, and the UK."

Official announcement: openai.com/index/introducing-dots

Best for: heavy ChatGPT users already on a Pro or Business plan.

Official video from the maker.

10. Claude Cowork

Ready-made Tier 2
Claude Cowork's real app, from Anthropic's Cowork page: a pricing-comparison task where Claude opens a browser and reads a pricing page

Real screen. Source: Anthropic's Cowork page.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

Cowork is Anthropic's agent for office work, with no terminal involved. Anthropic's own guide is direct about what that means: "Claude has access to the local files you grant it permission to access, and can take real actions on your behalf." It launched as a research preview in January, and on 16 September Anthropic merged Cowork and chat into one Claude, rolling out to Pro and Max plans.

Plugins bundle skills, connectors and sub-agents, and Anthropic says "Claude breaks complex work into smaller tasks and coordinates parallel workstreams to complete them." Tasks can run on a schedule, approval modes range from asking every time to acting on its own, and "Project sharing is available on Team and Enterprise plans", with view or edit access per person. Claude Pro costs $20 a month billed monthly. Of the six, it's the only one confirmed for the UK.

Official announcement: claude.com/blog/cowork-is-now-claude

Best for: office workers who want an agent on their own documents, and UK readers who can't get the others yet.

Official video from the maker.

11. Claude Projects

Build it yourself Tier 2 → 3
Claude Projects' real app, from Anthropic's official video: a project conversation beside its threads panel, with work waiting on the user, five threads working and two idle

Real screen. Source: Claude official video (0:19).

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

This is the redesign that started all the fuss. Announced on 17 September, it turns a project into a single long-running conversation: "Projects have threads that do the work and a coordinator that directs them." You brief the coordinator like a team lead, and it opens threads to do the work. Each thread is usually a session running on Anthropic's servers, working on its own copy of the code, and they keep going after you close your laptop.

Each cloud thread also starts with the project's knowledge already loaded: its written instructions, the CLAUDE.md files from its code (the instruction files Claude Code reads at the start of every session), their skills, and a shared MEMORY.md index of what the project has learned.

So it has orchestration and unattended work, the Tier 3 features. What it doesn't have yet is the reach. It's in a beta for Pro and Max plans, there are "no organization-level controls", and "A project belongs to one user. You can't share a project."

It's also where most of the excitement gets it wrong. The two features people are most excited about, a search layer over project files and team permissions, belong to the older chat Projects, not this redesign. We cover what actually shipped below.

Official docs: code.claude.com/docs/en/claude-projects

Best for: Claude Code users who want to hand off bigger jobs and let them run in parallel.

Official video from the maker.

12. OpenClaw

Build it yourself Tier 2 → 4
OpenClaw's real Control UI, from the OpenClaw blog: the Skill Workshop listing proposed skills, with a launch-checklist skill open for review and Apply, Revise and Reject buttons

Real screen. Source: OpenClaw blog.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

OpenClaw is the giant of the category, with 391,182 stars, an MIT licence and, in its own words, "no paid tier, hosted service, or token." Peter Steinberger built it as a personal assistant that runs on your own hardware and lives in your chat apps, with a personality file, Markdown memory, skills, scheduled automations and sub-agents.

It has also gone further than most toward teams. Its multi-user mode "lets several trusted people operate the same OpenClaw agent", with shared sessions "the whole team can open, steer, and take over" and named roles. That's the Tier 4 idea. But the same page says "A gateway is one trust domain." Everyone on the team is trusted with everything, which is fine for three co-founders and risky for a company. We come back to why in the security section.

Official project: github.com/openclaw/openclaw

Best for: a single power user, or a small team that fully trusts each other.

Official video from the maker.

13. Grok Bot

Ready-made Tier 2 → 4
A real Grok Bot chat from xAI's official video, with the bot replying 'Got it — watching the demo and turning it into a skill'

Real screen. Source: xAI official video (0:36).

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

xAI launched Grok Bot in beta on 11 August. It comes with SuperGrok ($30 a month) or a paid Cursor plan, and you reach it from desktop and mobile apps. Each named Bot keeps its own memory, files and preferences, can pick up a skill by watching you do a task once, and runs routines on a schedule or when something happens.

The catch is in xAI's own security documentation: "All of your Bots share one cloud computer assigned to your user account." The files, browser sessions and logins on it are available to every Bot, so separate Bots don't keep work apart.

Grok Bot is also the only ready-made option built to be shared. "A Team Bot is one Bot that your whole team talks to." Only its owner can change its setup, admins choose which groups can use it, and xAI says "A Team Bot never lends one person's access to another." Bots can also run in parallel and pass work to each other.

Official announcement: x.ai/news/introducing-grok-bot

Best for: power users and developers already paying for SuperGrok or Cursor.

Official video from the maker.

14. HAL

Build it yourself Tier 3

A simulated HAL session — illustration.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

HAL is the AI infrastructure I built to run StationX, and it's our worked example of Tier 3. It runs on Claude Code, but over the years it has been plugged into real systems: it manages cloud infrastructure across 6 providers and controls more than 108 tool integrations. Around 80% of a company serving half a million customers now runs through it. It scans our servers and code for vulnerabilities every week, ranks what it finds, reports into Slack, publishes content, and runs our sales and billing workflows.

Take one job it runs end to end. I point it at the open-source WordPress plugins that power much of the web, and it works through them hunting for security holes, across a pool of tens of thousands of plugins. When it finds one, it doesn't just flag it. It proves the hole is real by exploiting it on a safe test site, then produces a short video that shows the attack happening step by step. Finding the bug, proving it and packaging the proof used to be three separate jobs for a security researcher. HAL does the whole chain, and it has turned up real zero-day vulnerabilities: flaws the software's own makers didn't know were there.

It does the other side of the business too. Most of what you're reading right now came through HAL: the research behind our YouTube videos and articles, the thumbnails, the social posts, and the course and training material. One system runs the security work and the content work, because both are just workflows I've taught it.

The governance is what makes that safe enough to do. Hooks block dangerous commands before they run, every change to production needs a named approval, and a panel of AI reviewers from different model families checks code before it ships.

HAL is mine, which makes it a single-user system. It's the most powerful thing on this list for one person, and the next tier is about what happens when it isn't just one person. You can read how HAL works.

15. Paperclip

Build it yourself Tier 3 → 4
Paperclip's real dashboard, from the official README demo: agents enabled, tasks in progress, $309 monthly spend and a feed of agent activity

Real screen. Source: Paperclip README demo.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

Paperclip (MIT, 95,962 stars) sits a level above the personal assistants. Its README puts it neatly: "If OpenClaw is an employee, Paperclip is the company." It "orchestrates a team of AI agents to run a business": agents get job titles and a boss, recurring tasks run on schedules, and humans approve work with budget hard-stops.

Switch on its authenticated mode and it supports "Multiple Human Users" with roles and permissions, which is Tier 4 territory. Out of the box, though, it's "Optimized for single-operator local use." Its own memory layer is still on the roadmap.

Official project: github.com/paperclipai/paperclip

Best for: founders who want to run a company of agents with a human board on top.

16. Hosted HAL

Build it yourself Tier 4
Illustration of Hosted HAL: five team members in isolated workspaces all connected to one shared company brain, with a live usage HUD

Illustration.

At a glance

Identity & ContextMemoryCapabilitiesOrchestrationExecution & AutonomyInterface & AccessGovernance
Multi-userPer-user permissions

Hosted HAL is how we took HAL from one person to a team. It runs on our server, and each of our 20 to 30 users gets their own isolated workspace running Claude Code, with the shared StationX company context loaded by default: our standards, our commands, our skills and how we do things.

The shared part is curated. When someone builds a useful skill, it goes through a review and approval step and is then distributed to everyone. And the separation is deliberate: users can't see each other's data, and staff and customers see separate catalogues.

Content is the clearest example. A team member producing a video or an article doesn't need to read a document about how we do it, or ask the person who wrote the last one. They run a command, and the shared context already knows our research method, our thumbnail style and our house voice. The work compounds across the team instead of being trapped on one person's laptop.

Best for: this one isn't for sale. It's the shape we think every AI-native business will need.

ProductTierKindUsersPriceLicence
gstack1 add-onBuildOneFreeMIT
LifeOS (was PAI)2BuildOneFreeMIT
NanoClaw2BuildOne (roles)FreeMIT
Hermes Agent2BuildOneFreeMIT
Simon's Agentic OS2BuildOne operatorPaid membershipClosed
Gemini Spark2Ready-madeOneGoogle AI Pro, $19.99/moProprietary
Meta Muse2Ready-madeOneFree, or $20 / $100Proprietary
Cue2Ready-madeOneFree (invite)Proprietary
OpenAI dots2Ready-madeOneChatGPT Pro, from $100/moProprietary
Claude Cowork2Ready-madeOne (shared projects on Team)Claude Pro, $20/moProprietary
Claude Projects2 → 3BuildOneIn Pro / MaxProprietary
OpenClaw2 → 4BuildOne or small teamFreeMIT
Grok Bot2 → 4Ready-madeOne, or a team (Team Bots)SuperGrok, $30/moProprietary
HAL3BuildOneNot for salePrivate
Paperclip3 → 4BuildTeam (auth mode)Free self-hostMIT
Hosted HAL4Build20–30 teamNot for salePrivate

Which Kind of Assistant Is Right for You?

The ready-made assistants all do roughly the same things, so the choice comes down to what each one can reach and who is accountable when it reaches too far. Take a small-business owner who wants an agent to chase unpaid invoices. With Muse, Meta runs the computer, a separate Sentinel checks every request to the internet, and the agent asks before it sends anything. That's the easy route, and it's only available in the US and Canada. Run Hermes Agent on your own server and the invoices stay with you, apart from what you send to the AI model you connect. But if its approval prompts are switched off, nothing stops it emailing the wrong customer. Neither answer is wrong. They put the risk in different places.

For readers in the UK, the first question is simpler: dots, Gemini Spark and Muse aren't offered here yet, so Claude Cowork or a build-it-yourself option are the realistic starting points.

Work accounts are heading somewhere else again. Microsoft's Copilot Autopilot, previously called Scout, was due to move into private preview at the end of September, and Microsoft says it "lives in your tenant with its own identity, memory, computer and workspace": the company owns the agent, not you.

So instead of picking from a list of features, ask who runs the computer, who picks the model, whether you need a terminal, and who carries the blame when it goes wrong. The table puts the four kinds side by side.

KindExamplesYou getYou give up
Ready-made (no setup)Muse, Gemini Spark, dots, Grok Bot, Cue, Claude CoworkSign up and go. The vendor runs the computer, picks the model and sets the guardrailsControl. Your email, files and accounts sit on the vendor's computer, and many aren't in the UK yet
Managed by your employerMicrosoft Copilot Autopilot, Muse for Small BusinessAn agent your IT team switches on, with company controls and audit logsYou can't take it with you, and it only knows work
Hosted open sourceHermes Agent on Nous's cloudOpen-source software someone else runs, with the model of your choiceYou still own the permission settings, and pay for the model on top
Build or run it yourselfOpenClaw, Hermes Agent, NanoClaw, LifeOS, PaperclipFull control: your machine, your model, your rules, and it works anywhereYou do the security: network exposure, which skills you trust, approval settings

Back to the invoices: on Muse, a wrong email waits for your tap on Allow. On a self-hosted agent with approvals off, it has already gone.

Claude Projects: What Anthropic Actually Shipped

Because Projects got the most attention, it's worth being precise. This is what Anthropic's own documentation and announcement say, as of 2 October.

Claim you may have heardWhat the docs say
One conversation directs the workCorrect. A coordinator directs threads that do the work.
Threads keep running after you close your laptopCorrect for cloud threads. Local threads run only while that computer is awake.
Threads load your CLAUDE.md, skills and pluginsPartly. CLAUDE.md and skills load from the project's repositories. Plugins declared in a repository don't; you add them in project settings.
Projects search your files when they're too big for contextThat's the older chat Projects. The redesign's documentation doesn't mention it.
You can share projects with your team, with view or edit accessAlso the older Projects. The redesign is single-user and not yet on Team or Enterprise plans.

A few details nobody mentioned. There's no separate price, but it burns through your normal plan limits faster, because every running thread is a full session. There's a hard limit of 200 new threads a day. New projects default to Opus for everything. And idle threads that are watching a pull request (a proposed code change waiting for review) wake up, and spend again, when a test fails or a review comment lands.

How Claude Projects work: a coordinator briefed by a person opens threads, each a cloud session on its own copy of the code; threads can spawn sub-agents and all write to a shared MEMORY.md

The Best OpenClaw Alternatives, by What You Need

If you came here looking for OpenClaw alternatives, the right choice depends on what you want the assistant to do.

If you want…Look at
An assistant that learns and writes its own skillsHermes Agent
Strong containment, every agent in its own containerNanoClaw
A deeply personal assistant built around your life and goalsLifeOS
A finished system with support, for a non-technical business ownerSimon Scrapes' Agentic OS
Parallel work inside Claude Code, with nothing to installClaude Projects
A company of agents with human approval on topPaperclip
No setup at allMeta Muse, Gemini Spark or OpenAI dots (not in the UK yet), or Claude Cowork in the UK

From Personal AI to Company Brain: AI Is Going Multiplayer

Look at the list again and a pattern jumps out. Almost everything is single-user. Hermes calls itself "single-tenant". NanoClaw is "Built for the individual user". Anthropic's Projects belong to "one user". Even OpenClaw's team mode is "one trust domain". The ready-made assistants are mostly the same: Muse, Spark and Cue are built for one person. Grok Bot's Team Bots are the clearest ready-made exception.

The people funding the next wave of companies have noticed. Y Combinator's current Requests for Startups includes "Multiplayer AI", and the description could be a summary of this article: "AI agents are the most powerful new tool a team has, but it's the one thing people still use by themselves." Right now, it says, "working with AI is largely single-player", and the goal is to turn a team's agent work into "a shared, living thing instead of a thousand private threads."

The large vendors are already moving. OpenAI launched Frontier in February, "a new platform that helps enterprises build, deploy, and manage AI agents that can do real work." It connects a company's data warehouses, CRM and ticketing tools so that "It becomes a semantic layer for the enterprise that all AI coworkers can reference" (a semantic layer is one shared map of what the company's data means), and "Each AI coworker has its own identity, with explicit permissions and guardrails." At its developer day on 29 September, OpenAI added ChatGPT Space, where "teammates, ChatGPT, and your dot can build on shared knowledge".

PlatformWho it's forPrice
OpenAI FrontierLarge enterprises (first adopters include HP, Intuit, Oracle, State Farm, Thermo Fisher and Uber)Through sales only
Microsoft Agent 365Microsoft 365 companies managing their agents$15 per user per month, or in the $99 E7 bundle
Google Gemini EnterpriseGoogle Workspace companiesFrom $21 per seat per month
Claude EnterpriseTeams on Claude$20 per seat per month plus usage, billed annually, minimum 20 seats; strong admin controls, but the redesigned Projects with their shared memory aren't on Team or Enterprise plans yet

So the enterprise route to Tier 4 already exists, mostly aimed at large organisations and sold through sales teams. But you don't have to be one of them. The open-source route is taking shape for everyone else: Garry Tan's GBrain memory layer (30,490 stars) says "It works as a company brain too", and in company mode, "The whole team queries the same brain."

Illustration of GBrain: one shared brain queried by users with different permissions

Illustration.

This is the same shift we wrote about in AI-driven businesses. Y Combinator's Diana Hu put it bluntly: AI "should not be a tool your company just uses. It should be the operating system your company runs on." To do that, "you will need to make your entire company queryable." A company can't become AI-native on thirty private assistants that don't know what each other knows. It needs a shared brain, and that's Tier 4.

Single-player versus multiplayer AI: on the left, thirty separate private assistant icons not connected to each other; on the right, several people connected to one shared brain hub, each with a personal space

What Changes for Security When Agents Are Shared

This is the part most guides skip, and it's where StationX readers have an edge.

A single-user assistant has one trust boundary: you. If it reads a malicious web page or email and gets tricked by a hidden instruction (prompt injection), the damage is limited to what you can access. Share that assistant with a team and the boundary moves.

One trust domain means one blast radius. On OpenClaw's team gateway, anyone who can steer the agent can reach everything it can reach. A prompt injection through one person's inbox can act with the whole team's access.

Shared memory spreads bad data. NanoClaw warns that information "will cross-pollinate" between people in a group. If an attacker can plant a false "fact" in shared memory, every user's assistant now believes it.

Scoping is only as good as its weakest path. GBrain's company mode scopes access per user, but its own docs note that "OAuth source scoping only guards the HTTP MCP path." In plain terms, its permission checks cover one way in (the standard connector route) and not the others, so other paths to the same data aren't covered.

Autonomy multiplies cost and risk. Claude Projects threads run in auto mode by default and wake on their own when tests fail, and there are no organisation-level controls yet.

The ready-made assistants move the boundary in a different way. When the vendor runs the computer, you're trusting the vendor's design, and those designs are new.

Several agents can mean one computer. xAI's own documentation says all of your Grok Bots share one cloud computer, with the same files, browser sessions and logins, and warns: "Do not use separate Bots as a security boundary." Splitting work across bots doesn't keep it apart.

"Ask first" doesn't stop prompt injection. In January, while Cowork was still a research preview, security firm PromptArmor showed that a document with hidden instructions could make Claude Cowork upload a user's files to an attacker. In its words: "At no point in this process is human approval required." Approval prompts only cover the actions the vendor thought to guard.

Run it yourself and the hardening is yours. OpenClaw patched a one-click flaw that let an attacker steal its access token (CVE-2026-25253, rated 8.8 out of 10), and its own team admits that scanning skills "won't catch everything." Hermes Agent lets you switch its approval prompts off with one setting, leaving only a hard blocklist of catastrophic commands.

This is the part we took seriously when we built Hosted HAL as a shared system. Each person gets an isolated workspace, so users can't see each other's data, and staff and customers are kept to separate catalogues with no overlap. A shared brain doesn't have to mean a shared blast radius, but you only get that by designing the boundaries in.

None of this is a reason to avoid shared AI. It's the job description for the people who'll secure it: identity for every agent, per-user permissions, memory that records where each fact came from, and audit logs. The controls that keep an agent in bounds are the same ones we cover in our guides to AI guardrails and an AI governance framework. Frontier and Agent 365 are selling exactly those controls, which tells you where the work is going.

Which Tier Should You Build?

For most of us, the answer for the next few months is somewhere between Tier 2 and Tier 3. Build enough of a personal AI to feel what memory and skills do for you. Then start connecting it to the systems you actually work with, with proper guardrails, once you've felt the difference.

If you're not technical, a ready-made assistant is a fine first step, as long as it's offered where you live and you're careful about which accounts you connect. Start with low-stakes accounts, not your company email or your bank.

If you are technical, you don't need any of the products above to do that. Claude Code already gives you the harness, and memory, skills and hooks are files you can write today. The products are worth studying for their ideas: Hermes for memory, LifeOS for identity, NanoClaw for containment, Paperclip for running agents like a company.

Tier 4 is something you grow into once you have people to share it with. When you do, the companies that get there first won't be the ones with the best model. They'll be the ones whose whole team works from one brain, safely.

✅ The one thing to take away

You can now sign up for an agentic OS as easily as an email account, but the ready-made ones stop at Tier 2. The real value is the layer you build around a harness so it knows you, runs your work and, eventually, your team's. Start at Tier 2, climb to Tier 3, and treat shared Tier 4 as a design problem in security, not just a feature.

If you want to learn to build this properly — the four tiers, the memory and skills, the governance that makes it safe — that's exactly what we teach in the StationX AI Master's program. You don't start from a blank page, either: you get a copy of HAL's own skeleton, the config system that turns bare Claude Code into working infrastructure, and the packages that already solve the problems we've solved — to build on from day one rather than reinventing them.

Frequently Asked Questions

What is an agentic OS?

An AI assistant built on top of an AI agent tool, with a lasting memory, a profile of who you are and reusable skills, so it gets better at doing your work over time. It's also called personal AI, personal AI infrastructure or a second brain.

What's the difference between Meta Muse, OpenAI dots and Gemini Spark?

Very little in what they do. All three are ready-made personal agents that remember you, run on their own cloud computer around the clock, connect to your apps and ask before sending or paying. The differences are price and audience: Muse has a free tier and targets everyone, Gemini Spark comes with Google AI Pro, and a dot comes with ChatGPT Pro.

Can I use Muse, dots or Gemini Spark in the UK?

Not yet. As of October 2026, all three exclude the UK. Claude Cowork works in the UK, and the open-source options such as OpenClaw and Hermes Agent run anywhere because you host them yourself.

Is Claude Projects an agentic OS?

Partly. The 17 September 2026 redesign adds a coordinator that runs parallel threads with shared project memory, which are agentic OS features. It's currently single-user, in beta for Pro and Max plans, and has no organisation-level controls.

What is the best OpenClaw alternative?

It depends on what you need: Hermes Agent for self-improving memory, NanoClaw for container isolation, LifeOS for a personal life assistant, and Paperclip for running a team of agents. If you don't want to set anything up, the ready-made options are Meta Muse, Gemini Spark and OpenAI dots, or Claude Cowork in the UK.

Are OpenClaw and Hermes Agent free?

Yes. Both are open source under the MIT licence. You pay for the AI model you connect and for wherever you run it.

What is OpenAI Frontier?

OpenAI's enterprise platform for AI agents, launched in February 2026. It gives a company's agents shared business context, their own identities and permissions. It's available to a limited set of customers through sales.

What is a multiplayer AI agent?

An agent a whole team works with together, sharing context and handing work to each other, instead of each person using a private assistant. Y Combinator listed "Multiplayer AI" in its 2026 Requests for Startups.

About the Author

Nathan House

Nathan House, Founder & CEO of StationX

Nathan House has 30 years of hands-on cybersecurity experience and is Cambridge-educated, holding CISSP, CISA, CISM, OSCP, CEH, and SABSA. He founded StationX in 1999 — one of the UK’s first cybersecurity companies — and has secured £71 billion in UK mobile banking transactions and the London 2012 Olympics, advising clients including Microsoft, Cisco, BP, Vodafone, and VISA. He authored the world’s most popular cybersecurity course — a #1 Udemy bestseller taken by over 500,000 students — and was named Cyber Security Educator of the Year 2020, AI Security Educator of the Year, and a UK Top 25 Security Influencer 2025. A DEF CON speaker and featured expert on CNN, Fox News, NBC, and the BBC, Nathan leads StationX’s training of more than half a million students worldwide.