Skip to content
Back to Blog

AI Agents Ate My Boilerplate: What Actually Works in Enterprise Banking

Bashar AyyashJuly 9, 2026Updated September 21, 202611 min read2,112 words
AI Agents Ate My Boilerplate: What Actually Works in Enterprise Banking
TL;DR

Five AI agents run in production at a MENA bank, after four years working with it remotely. Each is a Laravel queue worker, not an API endpoint, so models swap without a mobile release. App Store submission time fell from 14 days to four. Every decision carries an audit trail, and none act without a human.

11 min read · 2,112 words

I work remotely with an enterprise bank in the region, and have done for four years. I do not name it here. Where a detail would describe its security posture rather than the engineering, I have generalised it.

Three agentic AI vendors pitched us last quarter. Every pitch had the same shape. A fast demo, perfect context, and a price that assumed compute was free.

Then you ask about throughput. About audit trails. About what a hallucinated compliance decision costs at two in the morning while the transaction pipeline is already saturated.

In a regulated system the useful question is not what the agent can do. It is what the agent is allowed to decide on its own, and the honest answer is almost nothing.

What are the five agents, and what is each one allowed to decide?

Five agents run in production. None of them decide anything by themselves. Each produces a draft, a score or a suggestion, and a person signs it.

AgentWhat it producesWhat it may never do
API contract enforcerLaravel request classes, TypeScript types, breaking-change PRsMerge anything touching a monetary amount
Compliance log archaeologistDraft regulator reports with an evidence chainFile a report
Mobile release checkA submission confidence score and failing screenshotsSubmit a build
Incident response shadowLikely cause, suggested rollback, prior incidentsExecute a remediation
Code review agentFindings on N+1 queries, authorization gaps, PII in logsApprove a pull request

That last column is the design. Everything else is implementation detail.

How do you generate API contracts from documentation that lies?

By treating the documentation as a hint and the response samples as the truth. We integrate with more than 40 third-party APIs: card networks, local switches, government identity services. Their documentation includes PDFs scanned in 2019, Swagger files that disagree with the endpoint, and services that return HTTP 200 with an HTML error page in the body.

The agent runs inside a GitHub Actions workflow rather than a chat window. It pulls the current spec or scrapes whatever passes for documentation, generates strict TypeScript from real response samples, writes Laravel FormRequest classes with validation extracted from the field descriptions, and opens a pull request whose description is the breaking-change analysis.

It still invents required fields. Anything touching a monetary amount carries a HUMAN_VERIFY label and does not merge without a person reading it.

The win was not speed. It was that the mobile team stopped shipping type errors into production.

What does a regulator's audit request actually require?

Ninety days of transaction history, delivered inside 48 hours, with an evidence chain a person can follow back to the query that produced each number. That used to take four engineers three days. It now takes one agent and one compliance officer.

The agent has read-only views onto our transaction warehouse, ClickHouse underneath, and retrieval over our internal compliance playbook, vectorised in Postgres with pgvector. The output is a Markdown report with the SQL embedded, so the officer can rerun any figure rather than trust it.

Embedding the query next to the answer is the part worth copying. A number a reviewer cannot reproduce is a number that fails the audit, whatever produced it.

Why does data residency force inference onto your own hardware?

Because the transaction data cannot leave. That single constraint decides the model, the hardware and the budget, before anyone discusses quality.

So we serve an open-weight model on hardware we own, with vLLM in front of it. It is slower than a hosted frontier model and it costs more GPUs. It is also the only version of this that a regulator will accept.

Latency makes the same argument from a different direction. Regional fibre to Frankfurt runs around 60 ms on a good day, before the model has read a token. Hosted AI services are priced and designed around a latency budget that assumes you are next door to the datacentre. For anything a customer waits on, local inference is not a preference.

Data residency is not a compliance checkbox you satisfy at the end. It is the first architectural constraint, and it eliminates most of the vendor shortlist before the evaluation starts.

How do you stop shipping App Store rejections?

By testing against your own rejection history rather than against a generic checklist. Guideline 4.2, Minimum Functionality, is subjective, and it is applied unevenly across reviewers in our region. A rejection costs a launch window, not an afternoon.

The agent runs Maestro and Detox flows in CI. The judgement sits on top: it reads previous rejection reasons from the App Store Connect API, generates cases that target those specific patterns, flags screens that read as a wrapped web page, and produces a submission confidence score.

MeasureBeforeAfter
Submission to approval14 days4 days

Apple did not get faster. We stopped submitting broken builds. The agent catches broken deep links, missing loading states and Arabic text truncating in right-to-left layouts. It cannot fix an app that is genuinely a wrapped website, and no agent can.

Should an agent remediate an incident on its own?

No. Not in a bank. Auto-remediation is a liability position, not an engineering one, so this agent is decision support for whoever is awake at three in the morning.

A PagerDuty webhook lands in a Laravel queue worker. The agent gathers recent deploys, database metrics and error logs, retrieves over our incident runbooks, and posts one Slack thread: likely root cause, a suggested rollback command, and the closest previous incidents.

During an annual bulk-payment run, card processing spiked 400%. The agent named the pattern correctly, bulk payment plus third-party processor throttling, and surfaced a manual failover procedure we had documented and never automated.

It also suggested a failover that would have broken our active-active architecture contract. A person caught it. That is the whole argument for keeping the last column of the table above empty.

What does a code review agent catch that generic tools miss?

The patterns that are only dangerous in your codebase. We tried the general-purpose review tools and they were too noisy to keep, because banking code has failure modes that a tool trained on public repositories does not recognise.

Ours takes PHPStan output as structured input and adds rules extracted from our own past incidents.

What it catchesWhy a generic tool misses it
Eloquent relationships loaded in a loop without with()It reads as valid code; only the query count shows it
Authorization checks bypassed through DB::raw()The gate exists, the query goes around it
Personally identifiable data in structured logsA logging line looks harmless out of context
Transaction boundaries spanning an external API callCorrect locally, wrong once the network is involved

Only the last two are really about banking. The first two are the same mistakes I write about in what yabasha.dev actually runs, at a scale where nobody gets fined for them.

What does "no code" actually cost by month six?

The pitch deck said no code. Here is the inventory six months in.

What "no code" turned out to beWhy it exists
Around 400 lines of YAMLGitHub Actions orchestration for every agent
Terraform for the GPU clustervLLM does not provision itself
Laravel queue workersModel calls time out, and a web request cannot wait
A custom evaluation frameworkEvery model swap has to be benchmarked before it ships

"No code" is a demo conceit. A production agent needs infrastructure, observability and fallbacks, which are the same disciplines any distributed system needs.

The timeout row is the one that bit hardest, and I wrote up the general version of that failure in building production AI agents that actually work. A queued job whose timeout is shorter than the model call it makes does not fail loudly. It writes nothing.

Where do these agents still get it wrong?

In four places, consistently. I list them because the failures shaped the design more than the successes did.

FailureWhat it looks likeThe guard
Invented required fieldsA generated request class rejects valid trafficHUMAN_VERIFY on anything monetary
Suggested remediation that breaks an architecture contractA plausible failover that violates active-activeNo agent executes anything
Schema mappings that are valid but business-wrongTwo fields that match on type and mean different thingsDomain constraints as guardrails, not prompt text
Confidence without calibrationA high submission score on a build that still failsThe score advises, a person decides

The third one is current work. I am testing whether agents can negotiate schema mappings across reconciliation vendors better than people do, on the theory that agents do not get bored and do not assume an equivalence is obvious. They converge fast. They also converge on mappings that pass every type check and are wrong about the business.

Why is every agent a queue worker rather than an API endpoint?

Because it lets us change the model without shipping a mobile release. Every agent in this stack is Laravel-side. The React Native app and the Next.js frontend call ordinary APIs, and the model is an implementation detail behind them rather than part of the interface.

Agent as queue workerAgent as API endpoint
Swap models without an app releaseModel choice is baked into the client contract
Add a rule-based fallback or a human step behind the same APIFallback logic has to live in the client
Scale GPU spend separately from user-facing computeThe two are coupled
Retries and timeouts are the queue's problemEvery caller reimplements them

This is the same shape I use at a far smaller scale on my own site, described in how I built an AI agent for my portfolio, and the cost argument for it is in cutting LLM API costs 60%.

Nobody using the app cares whether a model wrote the response. They care whether their salary arrived. The architecture follows from that, not from the model.

What is different about building this in MENA?

Three constraints that changed the architecture rather than the roadmap.

ConstraintWhat it forced
Connectivity that shifts under youRetry logic that would look paranoid in Silicon Valley: routing changes, filtering updates, seasonal traffic patterns that break a statistical model
Several regulators, different guidanceSAMA, the Central Bank of Jordan and CBUAE each publish their own AI governance direction, so we design for evidence extraction rather than accuracy alone
We cannot hire fifty ML engineersThe stack optimises for Laravel developers who can read a model card; the agent infrastructure is deliberately boring

The third constraint is the one that decided the most. A design only your two best people can operate is not a production design.

What would I check before trusting an agent in a regulated system?

Work from the audit backwards. If a reviewer cannot reproduce the number, nothing else on this list matters.

  • Every agent output a person signs is reproducible from the query or evidence embedded beside it.
  • No agent executes anything. It drafts, scores or suggests, and a person acts.
  • Anything touching money carries an explicit human-verification gate.
  • The data-residency constraint is settled before the model shortlist, not after.
  • Every model call runs behind a queue, so timeouts and retries belong to the queue rather than the caller.
  • Model swaps are benchmarked against your own evaluation set before they reach production.
  • The failure list is written down, and each entry has a named guard.

Trust but verify is the old banking phrase. For agents I would add: automate, but audit. Every decision an agent makes without a person seeing it is a decision you cannot defend.

Measured with

Figures as stated in the original post, written from four years working remotely with an enterprise bank in the region. The bank is not named, and operational detail that would describe its security posture has been generalised.

  • Laravel backend, Next.js web, React Native mobile
  • More than 40 third-party API integrations
  • Open-weight model served with vLLM on owned hardware; ClickHouse warehouse; Postgres with pgvector
  • Regulator audit shape: 90 days of history inside 48 hours
  • App Store submission to approval: 14 days before, 4 days after
  • Six months of production agents at the time of writing

Questions: @bayyash

Bashar Ayyash
AUTHOR

Bashar Ayyash (Yabasha)

AI Systems Architect for regulated industries — evals, harness design, AI security.

Bashar Ayyash is an AI engineer and dev lead in Amman, Jordan. 20 years shipping software, 4 years inside a regulated bank building production RAG and agent systems with evals, guardrails and monitoring — in Arabic and English. He writes at yabasha.dev and builds open-source tooling for AI-assisted development.

Newsletter

Practical AI + full-stack insights for MENA builders. No spam.

Read more on the blog

Browse the latest articles or explore the full archive.