A team with nobody dedicated to DevOps now hands the work to an AI agent and resolves the problem itself.
At a glance
| Before | Now | |
|---|---|---|
| Diagnosing a production failure | Not possible in-house, handed to someone outside | An AI agent diagnoses the servers connected to Alpacon |
| Infrastructure under management | Some servers connected | BNOW's whole estate connected to Alpacon |
| Outage to resolution | 4 hours or more at worst | Within an hour at worst |
What became possible once AI agents could reach production safely, through Alpacon.
BNOW
BNOW runs a solution that analyzes biometric data from animals. Capsule sensors and gateways collect it in the field; BNOW analyzes it and raises an alert when something looks wrong. Beyond Korea, it works with research institutes and food-security networks in Vietnam and Singapore, and runs field trials with farms there.
Challenge
An outage-sensitive service with nobody dedicated to infrastructure
BNOW takes data from field devices in real time. When a server stops, the service stops, and it shows immediately.
"If sensor data stops coming in, you know about it right away. We're a service that takes data in real time." (BNOW center director)
And nobody was dedicated to infrastructure. It is not that nobody handled it—the center director has run the servers and databases all along. What BNOW did not have is anyone whose only job is infrastructure.
So anything serious meant asking someone outside to fix it. And those things did not happen at convenient hours.
"Friday evening. Morning of a day off. It was a lot of those." (Elon Choo, CEO)
Alpacon was already connected, so the way into the infrastructure was open. What was missing was anyone inside the company who knew how to use it. Nobody could work out what had actually gone wrong; checking whether a server was still alive was about the extent of it.
"Before that we just used it to see whether something was dead or alive." (BNOW center director)
"We didn't know how to use Alpacon over here." (Elon Choo, CEO)
Since the data could not be allowed to stop, things did get fixed. The problem was never whether, but how long. Work that looks trivial in hindsight could run past four hours. Filing the request, working out the cause, and waiting on a reply each took their own slice of it.
AI agents looked like the answer, but could not touch production
When AI agents arrived, BNOW wanted to use them for exactly this: diagnosis and repair. What they could not do was point one at production, or let it carry out real work.
Hand an agent your server credentials and the work gets done. You also lose any boundary on how far it can reach. If it runs something irreversible, nothing stops it.
That matters more, not less, on a team with nobody dedicated to infrastructure: there is also nobody standing by to undo it.
Solution
Grant only the permissions, control what actually runs
Alpacon draws that boundary. An agent gets only the permissions a task requires, and the commands it actually runs are controlled. Instead of handing over credentials wholesale, you open a governed path.
That is the part BNOW noticed first—seeing Alpacon raise an alert and stop an agent that reached for something it should not.
"The one thing we could see was that when AI tried to connect to something indiscriminately, Alpacon threw an alert. That was great at first. 'Alpacon is blocking it.'" (Elon Choo, CEO)
Once it proved out, everything else got connected
At first only some servers were on Alpacon. Once the value was clear in use, the rest went on, Windows servers included.
Connecting was not the whole of it. The infrastructure was surveyed and written down. In the first session the database structure was analyzed and documented automatically, and the rule "back up the database and get user confirmation" was written into BNOW's own repository.
With everything connected, one agent could see the whole estate. Work that had been siloed per server came together, and from there real work could be handed over.
It took less than a week
BNOW was already evaluating AI agents when the Alpacon use case lined up with it. One demo, and they started applying it.
From that demo to the whole estate being connected took less than a week. No separate rollout period, no training program.
Results
Diagnosing everything from one machine
Finding out why an alert never went out used to mean going into each server and checking them one at a time. An agent had to be installed on every server, and each one instructed separately.
Now it happens from a single laptop.
"This alert didn't go out on such-and-such a day. Why not? Check three VMs, run a simulation on what they find, and tell me why." (BNOW center director)
Finding the cause, fixing the algorithm, pushing it to GitHub, and getting it onto the servers now runs as one motion. What could run past four hours is resolved within an hour at worst.
It gets handled outside working hours too
Outages still land on Friday evenings and weekends. What changed is what can be done about them at that hour.
When a call comes in from the field, the CEO now hears about it first and acts on it. The stretch spent filing a request and waiting for an answer is gone.
"Now I go into Alpacon and reboot it myself. The other way around from before. Back then I didn't know how." (Elon Choo, CEO)
Solved without hiring for it
BNOW needed someone dedicated to infrastructure and never filled the role. A company that had to ask someone outside for the cause of a failure now diagnoses and fixes it in-house.
"If you're a large company running a lot of servers with a server security team, it's a tool that team uses. For us, Alpacon becomes the security team." (Elon Choo, CEO)
