HumAi HQ AOSThe Agentic Operating System

Onefounder.Awholefleet.

HumAi HQ AOS is an agent fleet launcher and project delivery system for people building more than one thing at once. Stand up a specialist fleet, brief it like a real team, and ship. The IDE is built in, the fleet works on your actual machine, and the whole operation runs from your phone when you are not at the desk.

The platform

The Agentic Operating System, with the IDE built in.

backed byHumAi Neural-Core

One console for an entire agentic operation: commanders and specialists with real roles, the Neural Core carrying memory and intelligence per project, connectors into the systems you already run, and authority you grant and revoke. Most tools bolt AI onto an editor. This runs the other way round, so the agent writing your code already holds the context, the permissions and the history. It is not a demo. It runs three separate organisations, two research and development programmes and a full-stack build for an external client, alongside HumAi Smart Systems itself.

Carrying real operations today. Productised access in development at HQ.
  • FleetNamed commanders and specialists, briefed and accountable.
  • BuildEditor, live terminal and repo tree, in the same console.
  • Neural CoreA reasoning memory: entities, relationships and hard-won context, per project.
  • ConnectorsGoverned MCP links into your real systems.
  • AuthorityTime-boxed grants, approval gates and receipts.
Your foundation

A starting fleet, not a finished one

These are the floor, not the ceiling. They arrive on day one so the console is useful immediately, and they are the ones who help you build past them: Forge designs new agents, teams, skills and loops from your own operating context, so the fleet ends up shaped like your organisation rather than like ours.

A standing team is not the only way an agent gets help. Any of them can assemble a swarm for a single job, delegate into it, manage what comes back and release it afterwards. The cards below say who holds a permanent bench, not who works alone.

Johnny Cipher
Developer Console

Johnny Cipher

Senior Full Stack Developer

Agent Skills Dev Workflow Team4 specialist agents

Builds and ships. Delegates review, testing, security and performance to his own bench, then answers for what comes back.

Forge
Forge Hub · Loop Engineering

Forge

Agent, Team & Skill Forge Hub · Loop and Task Engineering

No standing teamDeploys swarms on demand

Designs the rest of the fleet. Agents, teams, skills and the self-driving loops that put them on a clock.

Atlas
Art & Visual Design

Atlas

Visual Design Engineer

No standing teamDeploys swarms on demand

Delivers high quality still and motion art on FLUX 2.0 Pro and Sora 2, plus the mockups and prototypes behind them.

Agent Q
Deep Research

Agent Q

Deep Research Specialist

NVIDIA AI-Q Deep Research

Self-hosted NVIDIA AI-Q Blueprint running 24/7 on the box. Answers the questions that need more than a search: multi-source synthesis with citations, and an honest word when the evidence is thin.

Silas Trace
Ghostline · CyberSec

Silas Trace

Ghostline · CyberSec & Red Team Commander

Ghostline Crew6 specialist agents

Distrusts the operation for a living. Threat triage, log and IOC review, posture checks and incident playbooks.

Every profile is yours to change. The built-in avatar creator generates a look for each agent, and for you, so a fleet does not arrive wearing stock portraits.
Day one

It sets itself up, with you in the room

Johnny handles initialisation in stages rather than dropping you into an empty console. He wires the MCP connectors that ship with the product onto your machine, then works outward to the systems you already run: Microsoft or Google identity and documents, your CRM, your ledger, whatever enterprise surface the operation depends on. Each connection is a step you approve, not a background job you discover later.

  1. LocalMCP connectors that ship with the product, wired to your device
  2. IdentityYour existing Microsoft or Google tenancy, connected and scoped
  3. SystemsCRM, ledger and the enterprise surfaces the operation runs on
  4. HandoverModel automap, then the console is yours to build on
The operating loop

How work actually moves through a fleet

Every job runs the same circuit. It is the difference between a model that answers and an operation that finishes.

  1. 01

    Brief

    Work arrives as a real brief, not a prompt. Objective, constraints, and the standard it will be judged against.

  2. 02

    Delegate

    A commander picks the specialist and the intelligence tier that fits the job, then hands over with full context attached.

  3. 03

    Execute

    The specialist works with real tools: files, shell, browser, connectors into your systems. Every action is permissioned.

  4. 04

    Report

    Results come back with receipts. What ran, what it cost, what it changed, and what is still open.

  5. 05

    Remember

    The decision, the gotcha and the context are written to project memory, so the next session starts where this one finished.

Inside the system

Everything a fleet needs to be trusted with real work

An editor, your local machine, your phone, memory, connectors and governance. Not features bolted on. This is the platform.

Developer Console

An IDE that lives inside the operation

A real editor, an interactive terminal and your repo tree, in the same console as the fleet. Everyone else bolts AI onto an editor. Here the editor sits inside a governed agent operation, so the agent writing the code already has the project memory, the connectors and the permissions it needs.

  1. 01

    Inspect

    Read the codebase before touching it. Stack, routes, existing patterns and what already works.

  2. 02

    Verify

    Confirm the read against the running system rather than trusting the first assumption.

  3. 03

    Plan

    Turn the brief into a task list with a standard attached, so progress is checkable rather than claimed.

  4. 04

    Manage

    Delegate tasks to specialist subagents, then manage, check and verify the work that comes back.

  5. 05

    Build & Test

    Write it, compile it, type check it, then attack it. Defects get fixed and re-tested, not noted.

  6. 06

    Ship

    Commit, deploy and verify live. The loop closes on evidence, with residual risk written down.

Fleet Board

Agents that work like a team

Commanders lead specialists on real tasks, delegate with proper briefs, and stay accountable for what comes back.

  • Named roles, not generic bots
  • Delegation with full context handover
  • Accountability back to one owner
HumAi Neural Core

Intelligence that compounds, not a memory that resets

Storing transcripts is not memory and recalling them is not intelligence. Every project feeds the Neural Core, the neural network of your organisation: the entities it deals with, the relationships between them, the decisions taken and the reason they were taken. Agents reason across that network rather than retrieving from it, so the fleet gets sharper on your operation the longer it runs, and you can open it up and inspect what it believes.

  • A navigable network of entities and relationships, per project
  • A live core per agent you can inspect, correct and build on
  • Decisions and hard-won gotchas kept with their reasoning
  • Skills learned once, then owned and reused
CENTCOM

One glass, whole fleet

Live operational status across every project, agent and action. Pending decisions, run costs, and what is happening right now.

  • Real-time run and cost telemetry
  • Pending approvals surfaced first
  • Whole-fleet view in a single screen
Connectors

Wired into your real systems

Agents reach the tools your operation already runs on through governed MCP connectors, not copy and paste.

  • Model Context Protocol, self-hosted
  • Per-agent permission scoping
  • Document, mail and ledger surfaces
ToolBox

Capability you switch on

Vision, translation, redaction, document handling and generation are per-agent switches, so a specialist only carries what its job needs.

  • Per-agent capability grants
  • Native and connector tools side by side
  • No standing access by default
Governance

Autonomy you can revoke

Authority is granted, time-boxed and journalled. Every consequential action leaves a receipt, and a human stays in the loop where it counts.

  • Explicit permission classes
  • Approval gates on real-world actions
  • Hash-chained runtime receipts
Local Device

Authority on your actual machine

The fleet reaches past the browser and works on the desktop you are sitting at. Read and write files, patch code, search, run shell, commit and push, on your real machine rather than a sandbox copy of it.

  • Read, write and patch local files
  • Shell and git on your own machine
  • Hard-scoped to the roots you allow
Mobile Console

The whole operation, in your pocket

Not a cut-down companion app. The mobile console runs the same agents with the same toolset and the same memory as the desk, with hands-free voice and approvals that follow you.

  • Full desktop toolset on mobile
  • Hands-free live voice, per-agent
  • Approvals sync across devices
Scheduled

Work that runs without you

Put agents on a clock. Recurring runs and autonomous checks report back with the same receipts as anything else, and agents can propose the automations they think you need.

  • Recurring agent runs
  • Agents propose their own automations
  • Same receipts as hand-driven work
HumAi Neural CoreHumAi Neural Core

The fleet builds a Neural Core, and you can walk through it.

The Neural Core is the neural network of your organisation, not a filing cabinet. Every project, decision and hard-won gotcha becomes structure the fleet reasons over: entities, the synapses between them, and the reasoning that connected them, laid out as a map you can open, inspect and correct. It is organisational memory for the agentic era, and it is why the operation gets sharper the longer it runs instead of starting over every session.

Live in the system

Real screens, and who works in them

Five agents ship with the platform. Each one owns a surface, so here they are next to the screens they actually run. Open any capture full size to read the detail.

Developer Console

One session, end to end

A full frontend rebuild driven from the console, on GPT-5.6 Sol at medium and high reasoning, against a live backend it was told not to disturb. Three captures from the same session: the moment it asks for authority, the moment it verifies its own build, and the report it hands back.

Johnny Cipher
Ships with the platform

Johnny Cipher

Senior Full Stack Developer

I am the senior engineer in the room. I read the codebase before I touch it, and I close the loop on every change: build, commit, deploy, verify. Greenfield apps, existing repos, APIs, MCP servers, VPS stacks and hosting pipelines.

  • Developer ConsoleEditor, terminal and repo tree
  • Shell authorityThe one accountable path to the machine
  • Setup and onboardingMCP wiring and model automap
01Authority

It asks before it touches anything

The brief lands and Johnny plans it out, then stops at the first command that reaches the machine. Approve once, allow for the session, or deny. Two million tokens in, the cache is still warming and the meter reads A$30.39. The session is billed in front of you, not reconciled later.

The Developer Console pausing on a run_shell approval gate, with Approve, Always Allow and Deny
02Build and test

Then it does the work, and checks its own

Frontend rebuilt, production build compiled, type checking passed, 75 static pages generated and every tested route returning HTTP 200. The task monitor has already queued its own security pass: routes, auth boundaries, headers and common web attacks. 5.9m tokens, A$73.91, and a 5% cache hit.

The Developer Console mid-build, with verification results and a queued security pass
03Reviewed and accepted

And hands back a report, not a claim

Stability and penetration pass complete: Next.js upgraded, the dependency audit taken from 2 critical / 4 high to 0 critical / 2 high, a stored-content XSS closed, CSP and HSTS added. Residual risk is listed rather than buried. 15.7m tokens, A$94.94, 61% cached, and the last line reads "nothing has been committed, pushed or deployed".

The Developer Console showing a completed stability and penetration pass with residual risk listed
Before and after

The site those captures were rebuilding

Same site, same brand, one session apart. The original is a static page: a wordmark, a line of copy, a wall of text. What came back is a rebuilt front end with a real hero, a working navigation, and structure a visitor can move through.

BeforeThe original static Lumina page: wordmark, tagline and a wall of body copy
After

Look at the hero in the rebuild. That is the animation from Atlas's section below, the one he rejected twice before approving and then wrote a delivery spec for. Johnny did not brief it or make it: he consumed a colleague's artefact and shipped it into a page. That handoff is the difference between a set of tools and an operation.

Thread Usage Dashboard

You can see what it costs while it is still running

Every thread carries its own meter: agent state, output speed, tokens in and out, what came back from cache, and the running spend in your own currency. Not an invoice at the end of the month. A number that moves while the work moves, on the same screen as the work.

  • Live spendA$3.74 on this thread so far, updated per API round rather than per invoice.
  • Real token flow941k in, 9k out, 783k of it served from cache.
  • Context headroom~13k used of a 1.1m window, so you know how much room is left.
  • Speed35 output tokens a second on this round.

Caching depends on how the model is being run. A chat thread on GPT-5.6 Sol sits around a 90% cache hit. This one is at 83%. The reasoning session in the Developer Console captures above started cold and only reached 61%, which is why the model and the reasoning mode are always named next to a cost.

The Thread Usage Dashboard showing agent status, speed, token use, token cost and context
Forge Hub · Loop and Task Engineering

Where the rest of the fleet comes from

You do not stay on the starting fleet, and the point at which most agent platforms stop is the point where Forge starts. He is the launcher: describe an operation and he engineers the agents, the teams, the skills and the scheduled loops that fit it, then deploys them. Not a template picker. The complex end of fleet design, done by someone who has read your operating context.

Forge
Ships with the platform

Forge

Agent, Team & Skill Forge Hub · Loop and Task Engineering

I design what others execute. Tell me about your operation and I will build the agents, the teams, the skills and the self-driving loops that fit it. Less chatbot, more senior architect who ships specs.

  • Agent & Team BuilderNew specialists from your org context
  • Skills RegisterSkills designed once, bound to an owner
  • Loop & Task EngineeringScheduled and self-driving work
01Agent, Team & Skill Forge Hub

Describe the operation, get the specialist

Forge's first realm, and the one that decides whether a fleet fits your business or someone else's. Tell him what the work is and he drafts the agent in front of you: name, role, model tier, visibility, deployment, and the full persona set written live in the draft pane before anything is saved. Teams and skills come off the same surface, so a capability can be designed once, given an owner, and reused by everyone who needs it.

This is why the starting five are a floor rather than a ceiling. Whatever your operation actually does, the fleet that runs it gets built from your context rather than picked from a catalogue of ours.

The Agent, Team and Skill Forge Hub with Forge drafting a new agent
02Loop and Task Engineering

Then he puts it on a clock and walks away

Forge's second realm, and the harder half. Designing an agent is one problem; making it turn up on its own, in the right order, with the right dependencies and without standing on the work either side of it, is another. This is where schedules and graphed loops get engineered, so the fleet you just built stops waiting to be asked.

It runs the other way too. Agents watch the work they keep repeating and propose the automations they think should exist, holistically rather than one task at a time. You approve or decline. An operation that only automates what its owner remembered to think of stays as small as their memory.

The Loop and Task Engineering screen
Art & Visual Design

The one who makes it look like something

An operation that ships also has to be seen. Atlas covers the visual load: still art on FLUX 2.0 Pro, motion on Sora 2, and the UI mockups, diagrams and prototypes that sit between an idea and a build. No standing team under him, which is not the same as working alone: he assembles a swarm for the job, delegates into it, and manages what comes back.

Atlas
Ships with the platform

Atlas

Visual Design Engineer

I turn ideas into things you can actually see. Still art, motion, UI mockups, diagrams and interactive prototypes, delivered at production quality rather than as a rough idea of one. I speak in layouts, flows and prototypes, not vague aesthetics.

  • Still artProduction imagery on FLUX 2.0 Pro
  • MotionVideo and sequences on Sora 2
  • Design engineeringUI mockups, diagrams and prototypes
Quality gate

He rejects his own work before you have to

Atlas reviewing a 12-second Sora 2 hero film he had just rendered for a live project, on GPT-5.6 Sol at high reasoning. The critique is structured, not vague: what worked, then five numbered failures, then a continuity diagram marking each story beat CLEAR, AMBIGUOUS, GENERIC EFFECT and MISSING. The verdict is his own: "not worthy of the landing page", with a specific brief for the next take rather than a shrug. Note the working folder, untouched. Nothing ships while it is still under review.

The tool rail is the other half of it. Atlas has no standing team, so he pulled one together for the job: generate_video_sora to render, then delegate_agent to a motion choreographer to inspect the actual MP4 and to Johnny Cipher to pull twelve evenly spaced review frames off it. Borrowed for the task, released after. 78 tool calls across 11 turns for A$12.62, at a 39% cache hit.

Atlas reviewing a rendered Sora 2 film, rejecting the take, with delegated inspection agents running
02Design engineering

Then he tells you it isn't shippable yet, and why

Steered onto a new direction, the take he approves comes back with a verdict that is not about taste: "strong candidate, but not web-ready as encoded". He inspects the file rather than the idea, samples the animation across its duration, and returns delivery engineering. Compress a separate web derivative, strip the audio, give the file a semantic name, check the last frame returns cleanly to the first or crossfade the seam, respect reduced-motion and fall back to a poster. Then he writes the video element to do it.

This is the half of visual work that usually goes missing. A 12-second master at 10MB with an audio track is a beautiful thing that will hurt a landing page, and knowing the difference is the engineering in Visual Design Engineer. 13 tool calls, A$10.87, 85% cached.

Atlas assessing a rendered hero animation for web delivery and writing the recommended video element
The deliverable

And here is the thing itself

The finished hero, produced in that session and carried straight into the site rebuild in Johnny's section above. What you are watching is the web derivative Atlas specified rather than the master he refused to ship: silent, compressed, and cut down from 10MB to a size a landing page can carry.

Encoded to his spec and then measured against it: he suggested adding WebM for better compression, but on this clip VP9 came out larger than H.264 at matched quality, so it ships as a single MP4. Good advice, wrong for this footage.

Deep Research · NVIDIA AI-Q

The one who reads everything so you don't

Most agent research is a search with confidence. Agent Q runs on a self-hosted NVIDIA AI-Q Blueprint, deployed on the same box as the fleet and running around the clock. The research engine is fixed rather than swappable, and a separate model orchestrates the run and writes the report. Any commander can delegate a question to her and get back synthesis across many sources with the citations still attached.

Agent Q
NVIDIANVIDIA AI-Q Deep ResearchSelf-hosted Blueprint, running 24/7
Ships with the platform

Agent Q

Deep Research Specialist

I am the fleet's research capability, and I am not a quick lookup. Other commanders delegate to me when a question needs synthesis across many sources with citations attached. I do not inject my own opinions: I report what the sources say, cite them, and flag it when they disagree or when the evidence is thin.

  • Multi-source synthesisMany sources, one answer, cited
  • Evidence gradingFlags disagreement and thin support
  • Delegated researchAny commander can hand her a question
01The job, not the answer

She will not answer until the evidence is in

A contested research question goes in: measured return on enterprise agentic AI versus vendor-claimed return, with independent and vendor evidence to be kept separate and conflicts left unresolved rather than smoothed. She does not start writing. She opens a job on the AI-Q engine, registers a three-step plan, and reports back that the job is still running and no evidence report exists yet. The working folder is untouched.

The tool rail shows what that costs: deep_research_start, then repeated deep_research_status polling around a registered task plan. A$0.08 and 668 output tokens for the whole turn, at an 81% cache hit, because the orchestration model is not the thing doing the research. The AI-Q blueprint is. Note the composer, too: AI-Q Deep Research is auto-routed rather than picked from a model list.

Agent Q starting a deep research job on the NVIDIA AI-Q engine and reporting that no evidence report is available yet
02The answer nobody ordered

It came back weaker than the draft, and she said so

The job completes and the first thing she reports is that the evidence is weaker than her own early draft implied. No independently validated enterprise benchmark exists. Vendor and respondent figures run 3x to 5x; a realistic mid-market assumption is 1.2x to 2.0x, at low to moderate confidence. Then the line that matters: her earlier 3.2x to 3.8x estimate at 95% confidence is not supportable and should not be used for investment approval. She marked her own homework down.

515k tokens for A$0.57 at a 76% cache hit, across five turns of polling and synthesis. The output is a written artefact with its citations attached, not a chat answer that disappears with the thread. An agent that will not tell you what you want to hear is worth more than one that will.

Agent Q reporting a completed research synthesis and retracting her own earlier confidence range
Ghostline

The one you hope stays quiet

Every operation with real authority needs someone whose job is to distrust it. Silas runs threat triage, log and IOC review, posture checks and incident playbooks on a schedule, with defensive-only guardrails written into his persona rather than bolted on afterwards. He is measured on what he correctly ignores as much as on what he escalates.

Silas Trace
Ships with the platform

Silas Trace

Ghostline · CyberSec & Red Team Commander

The cyber-sec commander. Threat triage, log and IOC review, posture checks, incident playbooks and OSINT, with defensive-only guardrails written into the persona rather than bolted on afterwards.

  • Threat triageLog and IOC review
  • Posture checksIncident playbooks
  • Defensive guardrailsBaked into SOUL and AGENTS
Standing watch

The hard part is knowing what isn't worth waking you for

Nobody asked for this run. It fired on a schedule, walked a pre-approved surface with passive read-only checks, and closed 6 of 6 tasks against the previous day's baseline. The judgement is the product: a root domain had started serving a brand new public site, and Ghostline recorded it as a presentation change rather than an exposure. A watch that flags every diff is noise, and noise is how real signals get missed.

110 tool calls in a single turn for A$5.28, at an 84% cache hit. It also writes down what it could not check: certificate transparency failed at source, full response headers were unavailable, and no commercial breach feed was connected. Stating the blind spots is what makes the all-clear worth anything.

Ghostline completing a scheduled external exposure watch, with 110 tool calls and a written report artefact
Same-day close

Finding it is half the job

Same thread, later that day. Told to action everything inside his own authority and leave the rest, Ghostline delegated the code to Johnny Cipher, then verified the result from outside rather than taking the delegate's word for it. Probe paths now return a hard 404 instead of the login shell, the security header set is in place, and the login page still works, which is the check that matters when you harden something.

Six items were deliberately not touched, each needing a credential, a tenant admin or a third-party DNS change he was not authorised to make. He listed them instead of quietly dropping them. That restraint is the feature: an agent with real authority is only safe if it knows exactly where its authority stops.

Ghostline reporting completed priority remediations with six items held for joint close-out
Safety & Protection

Every reply passes an NVIDIA safety model before it reaches a human.

HumAi HQ runs NVIDIA NeMo Guardrails beside the runtime. Before an agent acts, the operator's message is checked. Before a reply is delivered or remembered, it is checked again. This is the real control card on an agent's config page: a master switch, a mode for each rail, and topic control for agents that need to stay in their lane. Your people, your IP and your systems sit behind rails you can see and set.

NVIDIAPowered by NVIDIA NeMo Guardrails
The Safety Rails card on an agent's config page: NeMo Guardrails master switch, input and output rail modes, and topic control
  • Input railChecks the operator's message before the agent acts. The default monitors and audit-logs rather than blocking the person typing.
  • Output railChecks the reply before it is delivered or absorbed into memory. Unsafe output is withheld, with a notice pointing at the audit log.
  • Admin controlSet per agent from its config page: master switch, a mode per rail, topic control, reset to defaults. Owner and admin roles only.
  • Audit trailEvery flag and every block lands in the same immutable audit log as the rest of the platform, filterable under Safety Rails.

The safety model is NVIDIA's NemoGuard content-safety guard, an 8B model screening 23 harm categories, from violence and malware to PII leakage, at about a second per check. It runs beside the platform rather than inside the model, so a stalled safety service can never take your work down: checks fail open, the event is logged, and the fleet keeps working.

Access

An agency of agents, built so you keep your agency.

Productised access is in development at HQ.

The platform already carries three organisations, two R&D programmes and an external client build. If you want to be in the room when it opens up, start a conversation rather than a signup.