The problem with agents
Autonomous systems create operational risk.
A generative model can produce a bad answer. An agent can act on that answer: issue a refund, delete a record, expose a document or commit a workflow step. That makes permissions, approvals and recovery paths part of the core design.
The engineering work therefore extends beyond prompts. It includes permission boundaries, reversibility, escalation, logging and the ability to reconstruct what the system did and why.
- What is this agent permitted to do, and where does that permission come from?
- Which actions are reversible, and which need a human before they commit?
- How would you reconstruct a specific decision six months later, for a regulator or a court?
- What happens when a tool call fails, or returns something adversarial?
- Which of your AI systems are high-risk under the AI Act, and how is that reasoning recorded?
- Who is accountable when an autonomous system causes loss to a third party?
What we build
Agent systems, end to end.
Agent architecture
Single-agent and multi-agent designs: task decomposition, delegation between agents, tool and API surfaces, memory and state, retrieval, and the orchestration layer that decides what runs and in what order.
Guardrails & permissions
Scoped capability grants, least-privilege tool access, input and output filtering, spend and rate limits, blast-radius containment, and deterministic checks around non-deterministic components.
Human-in-the-loop
Escalation models that decide when a person is required, approval interfaces that make review realistic rather than theatrical, and the fallback behaviour when nobody responds.
Audit trails
Structured, tamper-evident logging of prompts, tool calls, decisions and outcomes, aligned with AI Act record-keeping expectations and useful during an incident.
Evaluation & red-teaming
Behavioural test suites, regression evaluation across model versions, adversarial testing including prompt injection through tool outputs, and robustness testing under degraded or hostile conditions.
Production operations
Monitoring, drift and cost observability, incident response procedures written for autonomous systems, and a rollback path that works when the failing component is a model.
Compliance
EU AI Act work for systems that take action.
We assess your systems, classify them and prepare the documentation needed to support that classification. The legal work is tied to the technical design because the controls depend on how the system actually behaves.
Assessment & classification
Inventory of your AI systems, classification under the AI Act's risk tiers, identification of your role as provider or deployer, and a gap analysis against the obligations that follow. Where the answer is uncertain, as it often is with agents, we document the reasoning and the assumptions behind it.
Documentation & governance
Technical documentation, risk and impact assessments, transparency notices, logging and human-oversight measures, data-governance records, AI governance frameworks and internal policies, and support through conformity assessment and regulatory enquiries.
How we run it
Five stages with concrete outputs.
Discovery
What the agent is meant to do, what it will be allowed to touch, and what the worst realistic outcome is. Output: a scoped design brief and a first risk view.
Classification
Where the system sits under the AI Act and adjacent law, and which obligations attach. Output: a written classification with reasoning and assumptions.
Design & build
Architecture, guardrails, permission model, logging and interfaces developed together. Output: a working system and its control design.
Evaluate
Behavioural, adversarial and regression testing against defined criteria. Output: an evaluation report you can hand to a security reviewer.
Operate
Monitoring, incident procedures, periodic re-evaluation as models and scope change. Output: a system that stays compliant after the launch date.
What you receive
Deliverables you can keep using.
Intellectual property in everything built specifically for you is assigned from creation and takes effect on payment. Where our own pre-existing tooling is embedded in a deliverable, you get a perpetual, irrevocable, worldwide licence to use, modify and maintain it. Third-party and open-source components are identified in the statement of work before work begins.
- Agent architecture and control design documentation
- AI Act classification memorandum with stated reasoning
- Risk assessment and, where required, impact assessment
- Technical documentation package and transparency notices
- Evaluation and red-team report with reproducible test suites
- Logging and audit-trail specification, implemented
- AI governance framework and internal policy set
- Incident-response and monitoring procedures for autonomous systems
Have an agent ready for production review?
Send us the prototype, the intended workflow and the concern blocking release. We will help define the controls needed to move forward.
