Skip to content
Technical brief · September 2026

AI Digital Humans: digital workers that use your tools the way a person does

A platform for hiring digital employees that sign in to the web tools a team already uses and do the work, with every action recorded, every risky step held for a person, and a way to get better from its own mistakes without ever moving a customer's data anywhere new.

aidigitalhuman.services AI Intelligent Services, a business name of 1001226048 Ontario Inc.
Co-authored with Claude Fable 5.1 by Anthropic Grok by xAI
1 · What it is

Hire a digital worker, not configure a bot

Most automation asks a business to build integrations: an API key here, a webhook there, a connector that breaks when a vendor renames a field. AI Digital Humans takes the other path. A digital worker logs in to the same web tools a human employee would, in an isolated cloud browser, and does the task by reading the screen and clicking, typing and checking, exactly as a person would. No integration project and no change to the customer's stack: onboarding is a login, and where a tool requires it, such as Gmail or Outlook, the worker signs in through that tool's official OAuth.

The product is framed around employment because that is the mental model that keeps everyone honest. A customer hires a worker at a level, assigns them a skill, watches them work live, reviews a replay, and can let them go. Nobody "creates an agent" or "runs an automation".

Levels
Entry, Mid and Expert workers, priced per worker per month. The level decides which tier of skills a worker may run; the platform refuses anything above it.
Skills
A catalog of typed skills in three tiers, from order status lookup and refund initiation to invoice processing and report distribution. Each one declares what it needs, the steps it takes, what counts as success and when to escalate.
Desks
A desk is a worker, or a small team of workers, packaged for one job, such as the Returns and "where is my order" desk. Desks are sold by the job they do rather than the level they are hired at.
Tools
Dozens of real products across support, commerce, CRM, finance, documents and email, plus any web app the customer adds themselves. If a person can drive it in a browser, so can a worker.
Work-Pilot
An assistant that answers questions about the workforce from the account's own records and turns a sentence such as "have the team clear this week's returns" into a planned mission.
2 · How a piece of work runs

Gates in, a hash chain through, a receipt out

Every run goes through one path, whether a customer clicked "assign", an email asked for it, or a desk picked it up. There is no second way to run work, which is what makes the guarantees below hold everywhere.

The launch gates

Before a worker starts, the platform checks that the account is in good standing, that a worker at this level may run a skill of this tier, that the runtime can genuinely drive this skill today, and that every tool the skill needs is connected. Each refusal names its own fix: "connect Zendesk first", "this is an Expert-level skill", "this skill is defined but not yet runnable". Nothing is refused with a generic error.

The recorded trail

Every action a worker takes is written by one recorder as a link in a hash chain: each entry carries the hash of the one before it, with a screenshot attached. Nothing writes an action by hand. Tampering with, deleting or reordering a step breaks the chain, and the replay shows exactly that.

Before any consequential step, a risk engine scores the action and explains its score in a sentence a reviewer can read. Some things hold on their own regardless of the score, such as crossing the account's money threshold, a bulk email, or a financial or legal field. A held run pauses cleanly, opens an approval request with full context, and waits. Approve resumes; deny aborts. Decisions are attributed and cannot be edited afterwards.

The receipt

When a run ends, it can export a work receipt: the full chain, a checksum of every screenshot, and a digital signature over the whole bundle. A public verifier page re-runs the chain arithmetic and checks the signature entirely in the visitor's browser. Nothing is uploaded and no account is needed, so a customer can hand a receipt to an auditor who trusts neither the customer nor the platform.

3 · A desk is a graph of workers

Several small workers on one workflow, connected only where data flows

A single worker running a skill is a straight line: step one, then two, then three. It works, and it is easy to replay. But the returns desk, the platform's flagship offer, is not a line. It is a graph: a few workers, each with one job, connected only by the data that passes between them.

Intake classify · order id ticket Chief policy pack · route kind = wismo kind = return WISMO Runner own session · own cookies Returns Runner own session · refund off same order? serialize draft Evidence re-reads the order evidence Checker read-only · own session holds A person decides approval inbox verdict decision Commit lanes open by risk result Receipts work + cost, always receipts Learn holds → proposals Blue boxes open a browser of their own. Nothing connects the two runners: they run at the same time unless they touch the same order.
The returns desk as it runs. Every arrow is named by the data that crosses it; where no data crosses, there is no arrow, and the two workers run concurrently. The Checker never receives anything from a runner except the draft itself, as data.

Two ideas carry this design.

An edge is data, or it is not an edge. The graph is validated when the code loads: every edge must name the data it carries, and the graph must have no cycles. Two nodes with no path between them are dispatched at the same time. The one hidden edge is a shared record: a reply and a return on the same order both write to that order, so units on the same order wait for each other, in arrival order, and units on different orders never wait at all. That rule is enforced both inside a process and across machines, so two runners in two places cannot open two returns on one order.

The checker shares nothing with the worker it checks. The Checker opens its own browser with its own cookie jar and a read-only credential, and reads the order again itself. Its context is built from those facts, never from the runner's transcript. Before every judgement it asserts, in code, that it holds zero runner messages and shares no session, cookie or credential with any runner. A worker cannot talk the checker into agreeing, because the checker never hears from it.

Who owns what

NodeSessionCookiesCredentialMay read the transcript of
Chiefnonenonenonenobody
Runnersownownthe tools they work inthemselves
Evidenceownownthe storenobody
Checkerownownread-onlyitself only
Commitownownthe tools it writes tonobody

The desk page in the product shows this matrix to the customer. It is not a setting; it is how the code is built.

4 · Loops live inside nodes

Produce, check, correct, repeat, and stop at the cap

A loop is the other half of the design, and it lives in exactly one place: inside a node. A runner does not produce a reply and hand it on. It produces a draft, runs a deterministic check against the order it just read, and if the check fails it corrects and tries again. The check is code, not another model's opinion: does the draft name the right order, is the sender the order's customer, does a refund exceed the order total, does the reply state the shipment status.

Read the order own session Draft reply · amount Check code, against the order pass On to Evidence draft is data fail → correct (a few times, each pass charged) past the cap Hold a person decides
The loop inside one node. A small fixed number of corrections is allowed; the failure past the cap becomes a hold rather than another attempt. Each pass costs a visit, so the cost receipt shows a runner that kept correcting, and a spend ceiling can stop the loop.

Two properties make these loops safe to run unattended.

  • Every loop has a budget. A correction cap, a spend ceiling per unit, and one escalation, the hold. When the ceiling is crossed, the unit holds and still writes its receipts, so a budget failure is never a silent one.
  • Recovery targets the failure class. A failure is classified before anything is decided about it. A read that timed out may be tried once more. A permission failure is never retried. A write is never retried, whatever its class. The same failure, unchanged, stops the loop. A person reading the trail sees "permission denied", not "failed" eleven times.

The Checker has a loop of its own: if its check fails, it re-reads the evidence in its own session and checks again, up to the same kind of cap. And the refund cap is applied there as well as at the Chief, because the Chief only ever sees the ticket, and the amount against the order total exists only once the draft and the evidence both do.

Off the runbook, a desk stops. Browser-driven work is only as good as the script it follows, so a desk never improvises. A ticket the policy pack has no worker for, such as a change of shipping address arriving at the returns desk, is held at the Chief for a person. It never reaches Commit and no tool is opened for it. A store page that does not match what the worker expects fails its check rather than being guessed at.

5 · Three worked examples

What a unit looks like from the outside

The names, orders and amounts below are fictional. They follow the shape of the platform's own demo tenant, not any customer's data.

Example A · a clean "where is my order"

Maya writes: "Order #1041 has said 'label created' for four days. Where is it?"

  • IntakeClassified as a where-is-my-order ticket. Order id 1041 pinned from the text.
  • ChiefPolicy pack applied: no always-hold conditions match. Routed to the WISMO Runner.
  • WISMO RunnerOpens its own browser, signs in, reads order 1041: In transit, placed 1 September. Drafts a reply naming the order and its status. Check passes on the first draft.
  • EvidenceOpens a second browser and reads the same order's customer, total, status and placed date, all four at once.
  • CheckerOpens a third browser with a read-only credential. Confirms it holds no runner messages. Judges the draft against the evidence: right order, right customer, status stated. Passes. Policy re-checked with the amounts known: nothing to hold.
  • CommitThe desk is in shadow mode: records what it would have sent, touches no tool.
  • ReceiptsWork receipt (what happened, under policy version 1, evidence hash) and cost receipt (cost by node, zero corrections) written.
  • LearnA clean run with zero corrections: the reference shape for this skill is refreshed.

Outcome: closed, in seconds, with three isolated browsers and nothing written outside the desk. In live mode the reply would have been sent and the receipt would say so.

Example B · a refund that needs a person

Jordan writes: "The whole of order #1204 needs to go back. Please refund in full."

  • Intake · ChiefClassified as a return, routed to the Returns Runner. The Chief sees only the ticket, so it cannot yet know the amount.
  • Returns RunnerReads order 1204: Delivered, total $920. Drafts a return with a $920 refund and a reply. Check passes.
  • Evidence · CheckerIndependent reads agree. The draft passes the Checker. Now the amount exists, the policy is re-evaluated: $920 is at or above the account's refund cap. Held: "Refund at or above your cap".
  • CommitRefused before anything else is considered: a refund at or above the cap with no recorded decision cannot commit, in any mode, whatever the status says.
  • ReceiptsBoth receipts written for the held unit, so the cost of getting this far is visible.
  • A personFinds "a $920 refund on order 1204 is waiting on you" in the review inbox, sees the draft and the evidence, and approves with a note.
  • Commit · ReceiptsThe unit resumes with the decision id attached. Receipts are rewritten over the old ones carrying the whole cost, and the trail shows who approved and when.

Outcome: the money lane opened for this one unit, by a named person. Denying instead would end the unit and teach the desk (see the next example).

Example C · the desk learns a rule, and a person arms it

Over one week, three return tickets arrive that mention a crushed box and a broken item.

  • Chief, three timesEach ticket matches the always-hold condition for damaged goods and waits for a person. Each hold becomes a candidate: a reason code, a unit id, a timestamp. Not the ticket text.
  • LearnThree similar candidates inside the window become one proposal in the review inbox: "Arm this rule on the returns desk? Always hold: damaged-goods claim." Proposing again does nothing; there is one question, not three.
  • A personApproves. The rule is armed with their name and the time, and the desk's policy version moves from 1 to 2.
  • Chief, next ticketHolds on the armed rule before any tool is touched. The receipt says "policy version 2", so anyone reading it later knows which rules judged it.
  • A daily jobCompares the wrong-order rate for units that ran before version 2 with those that ran after. If the rule made things worse, it is set back to "proposed" and the question returns to the inbox with the reason.

Outcome: the desk changed its own behaviour, but only through a person's decision, only in a way the Chief can actually enforce, and only while the numbers say it helps.

6 · How it learns and corrects

Holds become proposals; a person turns a proposal into a rule

A system that produces verdicts nobody acts on is writing reports. The desk closes that loop, but slowly and with a person in the middle, on purpose.

Hold or denial reason code · unit id candidate Similar, recently same key, one desk one proposal "Arm this rule?" review inbox a person Armed policy version + 1 applies Chief, next unit holds before any tool daily job: wrong-order rate rose after arming? back to proposed, ask again Nothing in this loop rewrites the policy on its own. The only write to the policy is the person's.
The learning loop. The system proposes; a person arms; the policy version moves so every receipt says which rules judged it; and a code job watches whether the rule helped.
  • Every hold and every denial becomes a candidate. A candidate is a reason code, a lane, a unit id and a timestamp. It is not the ticket, not the customer, not the reply.
  • Several similar candidates become one proposal, never two. Proposing is idempotent: a key that already has a proposed or armed rule gets nothing new.
  • Only a rule the Chief can act on is ever proposed. A hold on an outcome, such as "the checker failed", is not something the Chief can hold on before the work exists, so it is learned on the kind of ticket instead: repeated denied returns become "hold every return for a person?". Arming that genuinely changes the next run. The system never asks people to arm rules that would change nothing.
  • A person arms it. The arming records who and when, and bumps the desk's policy version. There is no automatic mode.
  • A daily code job judges the rule on the wrong-order rate before and after it. Too few finished units on either side is "not enough evidence", never "worse".

Two more ways it improves itself

  • Reference runs. When a person marks a replay as done right, the platform keeps the shape of that work: the order of actors and actions, the names of parameters and outputs, timing and score. Every later replay of that skill reports "matches reference" or "drifted" against it. Values are generalised before anything is stored, and a test proves that a known customer email written into a run never reaches the reference table.
  • The effort watchdog watches recent cost receipts per node. If a node's typical cost is over its budget, it steps that node's effort down and raises an alert. It never steps effort up. Raising it is a person's decision.

The principle underneath: a lesson moves down a ladder, from explanation to checklist to template to automated check to enforced policy. Holds are the raw material; proposals are the template; the armed rule is enforced policy; the daily job is the check on the check. At no point does a model decide what the policy is.

7 · Without seeing client data

Learning from shapes and codes, never from values

The learning described above works because of what it deliberately does not store. The distinction is between a customer's work product, which belongs to the customer's tenant and stays there, and the shape of the work, which is what the platform learns from.

Tenancy
Every route is account-scoped by construction. Cross-tenant access to a session or a unit answers "not yours to see" identically to "does not exist", so a probe learns nothing.
Operators
The platform's own operators see metadata and aggregates across accounts: how many runs, how many held, which tool connections are failing. They never see a run's parameters, results, screenshots or approval context.
Learning artifacts
Candidates carry a reason code and a unit id. Rules are sentences built from reason codes. Reference shapes store step names, outcomes and costs. None of these has a place for a customer's name, email, message or order value.
The Checker
Reads the evidence for itself, in the customer's own store, in a session nobody else touches. It judges facts it read against a draft it was handed as data. The judgement is deterministic code, so no customer content has to leave the tenant to be checked by anything.
Models
Model providers are pluggable and optional: Claude Fable 5.1 by Anthropic and Grok by xAI are the two the platform was built with. Questions such as "what needs me?" are answered by a deterministic engine straight from the account's own records, and every model path degrades to that engine. A model is never asked to restate the database.
Failure text
Errors, stacks and failure notes are redacted on the way in: secrets by key name and by value, plus anything a browser may have echoed back from a form.
Secrets and screenshots
Credentials are encrypted at rest with authenticated encryption; screenshots live in private storage; every customer-supplied web address is checked against internal and private network ranges before the server ever fetches it, and re-checked on every redirect.
Receipts
A work receipt carries hashes, not content: the chain, a checksum per screenshot, and a signature. The public verifier needs no upload, because the customer keeps the receipt and the verification runs in their own browser. A desk's receipt also lists every rule it applied with its verdict, the evidence it judged against, and how to challenge the outcome.
Who owns a mistake
The customer remains the operator of record. Nothing customer-visible, money-moving or production-changing happens without either a rule the customer set or a decision a named person recorded by id. One refund per order, whatever the retries: a commit is idempotent per order and action, so a crash after the write cannot move the money twice.
Deletion
Everything hangs off the account. Deleting an account removes its workers, sessions, actions, units, candidates and rules. The test suites prove it by deleting the throwaway accounts they create and checking nothing is left.

What "the system learns" actually means here: it counts how often a kind of hold recurs, it remembers the shape of a run a person approved, and it watches its own cost per node. All three are done with identifiers, codes and structure. The customer's tickets, orders and replies stay in the customer's account, behind the customer's access control, and are read only by the worker doing the customer's own work.

8 · How fast

Measured, then made faster: where the time goes before a ticket is complete

Speed is a property of the system, not a claim, so it is measured the way everything else here is: from the trail. A report reads the duration of every node in every stored desk unit and the gap between consecutive actions in every live browser session, then runs fresh units against a real store with browser warm-up off and on, timing every launch, sign-in, navigation and read. The numbers below are one afternoon's run of that report, from a laptop in Ontario against a store hosted on Vercel. On the deployment itself, next to the database, every round trip is a fraction of these.

Per action, live browser session
1.79s 1.13s
37% faster
before 1.79s after 1.13s

Median gap from one recorded action to the next: a click, a keystroke, a navigation, a page read.

One whole skill run
23.2s 13.2s
43% faster
before 23.2s after 13.2s

Order status lookup and update, first recorded action to last. Same skill, same machine, same store.

One desk unit, live reads
2.32s 1.27s
45% faster
before 2.32s after 1.27s

A where-is-my-order ticket through Intake, Chief, Runner, Evidence and Checker, three isolated browsers.

What the measurement found

  • The per-action cost was not the browser. Every recorded step was a chain of round trips: a check that a person had not pressed Stop, then a screenshot, then the screenshot upload, then the row in the hash chain, then the chain head, then a fixed pause so the live view reads as work rather than a flash. About 1.9 seconds per action, whatever the action.
  • A desk unit opened its three browsers one after another, and signing each one in was the single largest phase.
  • In shadow units the bookkeeping nodes were slowest, waiting on one database write after another.

What changed, and what did not

  • The stop check and the screenshot now run together; the upload runs beside the row insert, which is safe because the chain hashes the screenshot's key, chosen up front, not its bytes; and the pacing pause overlaps the recording as a floor between actions instead of a tax after it. The check still happens before every action, and the chain is re-verified end to end.
  • The desk starts launching and signing in the Evidence and Checker browsers the moment the runner is chosen, under the runner's work. Each warmed browser is adopted by exactly one session of its node, with its own cookie jar and its own sign-in. The Checker still shares nothing with the worker it checks.
  • Learn writes its candidates and proposals at once.
  • What did not change: every action is still recorded before the next begins, every hold still waits for a person, and sign-in per browser, about two thirds of a second, stays. That is the price of isolation and it is paid on purpose.
Per browser phaseMedianWhat it is
Launch0.17sA fresh browser with its own cookie jar
Sign in0.67sThe store's own login, per browser, never shared
Navigate0.21sOpen the order page
Parse0.04sRead customer, total, status and placed date, all four at once

Why publish this: a vendor who measures where its own time goes, in the same trail it hands its customers, can be held to it. The report ships with the platform and is re-run after every change to the runner. Start a 30-day pilot and watch a unit walk the desk in your own tools, or read what the product is at aidigitalhuman.services.

9 · What is new here

Where this differs from a bot with a prompt

None of the individual ingredients is exotic. What is unusual is that they are combined into one product with one path for all work, and that the safety properties are enforced by structure rather than by instructions a model is asked to remember.

  1. No integration surface at all. A worker uses the customer's tools through a browser, so onboarding is a login, not a project. The customer's stack does not change, and the customer's data is not copied into a new system to make the automation possible.
  2. Every run is evidence. A hash-chained, screenshot-backed, signed trail that a third party can verify offline turns "the bot did it" into something an auditor can check. Most automation products keep logs; few produce a receipt that survives distrust of the vendor.
  3. A desk is a graph, and the checker is adversarial by construction. Splitting a workflow across small workers with named data edges is what makes concurrency safe and the checker's isolation provable. Asking one model to double-check itself is a confidence loop; a checker on its own session with its own read of the evidence is verification.
  4. Lanes open by blast radius, not by confidence. Reads are open. Sending to a customer, moving money and writing to production stay closed until a person or a deterministic gate opens them, per unit. The whole graph can run in shadow mode, writing both receipts and touching nothing outside the desk. Going live is a switch a person flips.
  5. Learning that cannot rewrite its own policy. The system can only propose. A person arms. A code job can only demote. And what it learns from is shapes and codes, so the learning layer holds no customer values to leak.
  6. The cost of a loop is on the receipt. Every pass round a correction loop is charged, every unit writes a cost receipt beside its work receipt, and a routine with a finished unit lacking a cost receipt disables itself. Spend cannot become invisible.
  7. Everything degrades to something deterministic. Without a model, the assistant still answers from the records, the desk still runs its graph, and the demo still walks end to end. A model adds judgement where judgement is needed; it is never load-bearing for safety.

See it for yourself: the product behind this brief is at aidigitalhuman.services, and every pilot starts with a real unit walking the desk in your own tools. Proven, not promised: a regression net covering the gates, the chain, the risk engine, the desk's rules and the demo script runs before every release and again against the live deployment afterwards. A fix that only holds in a development build is not treated as a fix.

10 · Glossary

The words the product uses

Digital worker
An employee-shaped unit of capacity: a name, a level, assigned skills, a mailbox, a history.
Skill
A typed definition of one job: parameters, steps, success criteria, escalation rules, required tools.
Desk
One or more workers packaged for one workflow and sold by the job. The returns desk is a graph of several.
Policy pack
The rules a desk runs under: a refund cap, a return window, always-hold conditions, and every rule a person has armed. Versioned.
Isolated session
A browser with its own cookie jar and credential, opened for one node of one unit and closed after.
Approval inbox
Where held work waits for a person. Risk holds, desk holds, drafts, suggested updates and rule proposals all land there.
Work receipt
What happened, under which policy version, with whose decision, plus a hash of the evidence. Signed, portable, publicly verifiable.
Cost receipt
What a unit cost, by node, with the number of corrections. Every finished unit writes one.
Replay
A scrubbable playback of a run's trail with its screenshots, and its verdict against the reference run.
Shadow mode
The whole graph runs and records; nothing outside the desk is written.

← aidigitalhuman.services · The Returns & WISMO Desk · Start a 30-day pilot

AI Intelligent Services, a business name of 1001226048 Ontario Inc.

Co-authors: Claude Fable 5.1 by Anthropic and Grok by xAI. The platform was designed, written and verified in collaboration with both, and its model layer is pluggable so either can sit behind Work-Pilot and the workers.

This brief describes the design at the level of ideas and behaviour. Implementation details, thresholds and internals are intentionally omitted.