How BitSafe Runs, Companion Essay
The hardest question in our AI program was also the most reasonable: is this system worth what we spend on it?
We wanted a clean answer. Hours saved. Revenue caused. Headcount avoided. A single return figure that could sit on a leadership scorecard.
We could not produce one we trusted.
The problem was not a lack of positive examples. Routine work completed more consistently. Information became easier to retrieve. Some workflows moved with less manual coordination. The problem was causality. We could see changes, but we could not isolate what would have happened without the AI system.
So we stopped trying to compress the whole program into one ROI number. We now use three measurement layers: adoption and maturity, operational reliability, and cost or usage. Each layer answers a different question. None of them proves business impact on its own.
The discipline is to label inputs, proxies, and outcomes honestly.
Why the obvious ROI calculation breaks
A credible return calculation needs a defensible counterfactual.
If a person uses AI to finish a report, how long would the same report have taken under the same conditions without it? If a proactive agent catches a stale record, would a colleague have noticed later? If a workflow keeps a follow-up from slipping, what is the value of an error that never occurred?
Real company work does not provide controlled duplicates. The task, person, timing, and surrounding process all change at once.
Attribution creates another problem. A sales process may improve because the data model changed, ownership became clearer, and automation removed a delay. Assigning the gain to AI alone would ignore the operating work around it.
This is why precise “hours saved” claims often rest on estimates presented as measurements. The same weakness applies to revenue attribution and headcount avoidance. Without a sound comparison, precision creates confidence without evidence.
We refuse to make those claims.
Layer one: adoption and maturity
The first layer asks whether the system is being used and whether use is becoming part of real work.
Adoption metrics are inputs. They can include active users, completed sessions, recurring workflow use, or the share of a team that interacts with the system during a defined period.
These numbers matter because an unused system cannot create value. They also have strict limits. High activity may reflect confusion, repeated attempts, or novelty. A person can generate many sessions without changing a single workflow.
That is why usage should be paired with maturity.
A maturity measure asks how deeply AI is integrated into the way someone works. An illustrative ladder can remain simple:
Aware: understands what the tools can do.
Trying: uses them for occasional tasks.
Using: relies on them in a recurring workflow.
Integrated: has redesigned part of the workflow around them.
Transforming: can show that the operating model itself changed.
The exact labels matter less than the evidence required to move between them. “Integrated” should mean more than enthusiasm. A person should be able to name the workflow, the old method, the new method, and the failure mode if the tool disappeared.
Usage is an input. Maturity is a proxy for process change. Neither is an outcome such as revenue, customer retention, or product velocity.
Layer two: operational reliability
The second layer asks whether the system completes the work it was assigned.
This is where an internal AI stack starts to look less like a software subscription and more like an operating service.
A useful reliability scorecard can track:
How many expected jobs started and completed.
How many required a retry or human rescue.
How many ended without a recorded outcome.
Whether the result reached the intended system of record.
Whether an alert led to a clear next action.
Whether the work met its review or accuracy standard.
These are work-completion measures. They tell us whether the engine is dependable inside its defined scope.
They still need careful definitions. A task that produces a document may be complete from the system’s perspective but unusable after review. A monitor that fires on time may be operationally healthy while creating too much noise. A workflow can meet its completion target and still solve the wrong problem.
Quality therefore belongs next to completion. We prefer narrow checks that match the job. A research task should preserve sources. A data update should land in the correct record. A public draft should pass the relevant review. A monitor should remain quiet when nothing requires action.
These metrics are proxies for operational value. They do not prove broader business outcomes, but they show whether the system is earning the right to stay in the workflow.
Layer three: cost and usage
The third layer asks what the system consumes and where that consumption comes from.
Cost sounds objective, but AI cost telemetry is rarely complete on day one. Different execution surfaces can record usage differently. Some work may be easy to attribute to a task while other work sits in a shared account. A model call may be visible even when the surrounding engineering and review cost is not.
We therefore treat cost as a measured input with a coverage statement, not as the one number that always holds still.
The practical questions are:
Can we attribute meaningful usage to a workflow or class of work?
Can we see when the model, context size, retry pattern, or schedule changes?
Can we compare the cost of repeated work over time?
Can we set limits that stop a small error from becoming a large bill?
Can we explain what the telemetry does not cover?
A bounded cost is useful even when the benefit is not reducible to one figure. It lets leadership decide how much uncertainty the program can carry while the evidence improves.
Usage also helps diagnose reliability. A sudden cost increase without a corresponding rise in completed work is an operating signal. It may indicate a routing mistake, repeated retries, oversized context, or a loop that stopped making progress.
Keep the categories separate
The easiest way to misread an AI scorecard is to let one category stand in for another.
More sessions do not prove more value. A higher maturity score does not prove revenue. A strong completion rate does not prove that the selected work matters. Lower cost does not prove that quality held.
Read the layers together.
Adoption tells us whether people are engaging with the system. Maturity shows whether use is becoming part of repeatable work. Reliability shows whether those workflows complete to a defined standard. Cost shows whether the operating burden remains visible and bounded.
Business outcomes sit outside that chain and require their own evidence. If a team wants to connect an AI workflow to an outcome, it should state the hypothesis in advance, define the expected mechanism, and identify competing explanations.
That will not always produce a causal answer. It will produce a more honest one.
What belongs on the scorecard
A practical leadership view can stay compact.
Adoption input: a clearly defined usage measure over a fixed period.
Maturity proxy: the share of relevant people with AI embedded in a recurring workflow.
Reliability proxy: completion, review pass rate, or silent-failure rate for selected workflows.
Cost input: usage and spend with attribution coverage and limits.
Outcome hypothesis: one business result the workflow is expected to influence, without claiming causality before evidence exists.
The scorecard should also show definitions. A number that cannot be reproduced from its rule will become a debate about interpretation.
We have not reached a final measurement system. Cost attribution is improving. Reliability definitions need to match each workflow. Maturity still contains self-reporting. Those limits belong in the operating conversation rather than in a footnote nobody reads.
The operating lesson
Measure the system you actually have.
Track use without calling it value. Track completion without calling it ROI. Track cost without claiming complete coverage. Treat outcomes as outcomes and demand a defensible link before crediting AI.
This approach gives up the satisfying headline. It gains a scorecard that can survive scrutiny.
We still cannot prove exactly how much return our AI system creates. We can see whether people use it, whether workflows depend on it, whether assigned work completes, and whether cost remains controlled. That is enough to make better operating decisions while the evidence matures.
Continue the series
Relevant core article: We Automated the Automator. Here’s What Still Needs a Human.
Series hub: How BitSafe Runs on AI
Subscribe to the BitSafe newsletter for more practical notes on operating AI systems.

