Skip to Content

When Your Company's IT Lives in One Person's Head

Key person risk in IT is not an HR topic: it is a continuity risk, and it carries a price.
August 10, 2026 by
When Your Company's IT Lives in One Person's Head
Kleber Leal by Zamak Portal

On an accounting close Friday, the billing system at a distributor with 380 employees stopped responding. The team did what seemed obvious: they restarted the server. The server never came back. The only person who knew that this environment required a specific startup sequence, and who kept the administrative credential in a personal password manager, was aboard an international flight for the next nine hours. Monthly billing was frozen for a day and a half. No equipment had failed.

Episodes like this rarely make it into the incident report for what they actually are. They get logged as a system failure, when the honest diagnosis would be something else: the company was running a critical process whose execution depended on a single person's mind. That is a real operational liability. It doesn't show up on the balance sheet, it carries no accounting provision, it isn't audited by anyone, but it precisely defines how long the business stops when that person leaves, gets sick, takes vacation, or is simply out of reach.

The natural temptation is to treat the issue as a people question: retain them better, pay more, hire a backup. That's a framing error. Knowledge concentration is a governance and continuity question, and it's solved the way any governance risk is solved: by making visible what is implicit, turning informal practice into a verifiable process, and creating redundancy where today there is a single line of defense.

The liability that doesn't show up on the balance sheet

Dependence on one person rarely comes from negligence. It comes from competence under pressure. Someone solves an urgent problem with a clever tweak, the tweak works, and what was an exception becomes a silent standard. Repeat that over five or seven years and the organization accumulates hundreds of technical decisions that only make sense to the person who made them. Engineers call this the bus factor (bus factor, the number of people who would have to become unavailable for a system to stop being operable). At a good share of midsize companies, that number is one.

Labor market pressure makes the picture worse. CompTIA, in its State of the Tech Workforce 2025, projects that technology occupations should grow at a rate close to double the average for the economy as a whole over the next decade, which means permanent competition for experienced professionals. ISACA, in its State of Cybersecurity 2024, found that 57% of organizations operate with security teams below the level they need, and that a significant share take more than six months to fill a senior technical position. Translated into business language: when the person who holds up the operation leaves, the exposure window isn't weeks, it's quarters.

The effect shows up first in time. MTTR (Mean Time To Repair, the average time to restore a service) is the metric that connects infrastructure to results, because every minute of it is a minute of stopped operations. When the procedure is documented, the clock starts when someone opens the document. When the procedure lives in one person's memory, the clock starts when someone manages to reach them. That difference, invisible on the org chart, is usually what separates a forty-minute incident from a two-day incident.

Nor does the cost end when the system comes back. There's the rework of reprocessing orders, redoing reconciliations, and answering customers. There are senior managers' hours diverted to managing a crisis instead of generating revenue. There's the reputational damage, which sends no invoice but has a long memory, especially in sectors where the customer has an alternative a click away. Putting a value on this is a simple and revealing exercise: a downtime cost calculator settles in minutes what intuition underestimates for years.

There is also the dimension that only surfaces when the company changes hands. In acquisition, fundraising, or major contract renewal processes, due diligence asks whether the operation is institutional or personal. An outdated inventory, the absence of written procedures, shared administrative credentials, and privileged access without traceability are treated as priceable risk, and priceable risk becomes a discount on the valuation or a holdback clause. Gartner, in mapping the priorities of infrastructure and operations leaders for 2025, places the skills gap and the lack of standardization among the central obstacles to execution, precisely because a non-standardized environment neither scales nor transfers.

How to convert personal dependence into institutional capability

The first move is to change the question. Instead of how do we retain this person, ask how do we make this knowledge belong to the company. Start with a map of critical processes: list the twenty procedures that, if stopped, stop the business, and write next to each one the name of whoever can execute it without consulting anyone. Wherever there is only one name, there is an identified risk. This exercise usually takes an afternoon and produces more strategic clarity than quarters of debate about the technology budget.

The second move is to build three assets. The first is the set of runbooks (step-by-step guides for executing and recovering each critical process, written to be followed by people who didn't create them). The second is a living inventory of systems, contracts, vendors, and maintenance windows, updated by routine rather than by memory. The third is privileged access management, or PAM (Privileged Access Management, control over who holds administrative credentials), with a corporate vault, periodic password rotation, and immediate revocation upon offboarding. None of the three is a technology project. All three are governance projects.

The third move is to solve coverage outside business hours, which is where personal dependence hurts most. A structure of managed IT operates a NOC (Network Operations Center, an operations center that monitors availability and performance) and a SOC (Security Operations Center, a center that monitors and responds to threats) on a 24x7 basis, with defined escalation. The sensitive point deserves to be spelled out: co-management does not replace the internal team. It takes the internal team off perpetual on-call duty and returns those people to the work that generates value, which is understanding the business and improving processes, not answering an alarm at three in the morning.

The fourth move is knowing how to buy. Demand an explicit responsibility matrix in the RACI format (who executes, who approves, who is consulted, who is informed), a service level agreement by severity with the clock starting at detection and not at the ticket, contractual ownership of the documentation produced, and monthly reports readable by non-technical people. Before signing, ask for a real test of nighttime escalation. A vendor that will not agree to be tested before the contract will not respond well during it either.

Five questions every manager should ask

  1. How many critical processes today depend on a single person to be carried out?
  2. What is the real cost of a prolonged outage because the person who knew how to fix it was not available?
  3. What evidence does a buyer or investor look for to assess whether the IT operation is institutional or personal?
  4. Where does the internal team's responsibility end and the managed back-office support begin?
  5. How do you turn tacit knowledge into a company asset without paralyzing the operation during the transition?

How many critical processes today depend on a single person to be carried out?

The measurement is objective and requires no tool. For each critical process, record how many people can carry it out in full without consulting anyone else, and how many have actually done so in the last twelve months. Knowledge that exists only on paper, without recent practice, counts as half a point. The concentration indicator is the proportion of processes with coverage equal to one. Above 30%, the company does not have an operation, it has an arrangement.

The managerial value of that number lies in making it trackable. Put it in the same report that already shows delinquency, inventory turnover, or margin by line. An indicator that shows up in the board meeting gets budget and a deadline. A risk that only shows up when it becomes a crisis gets only people to blame.

What is the real cost of a prolonged outage because the person who knew how to fix it was not available?

The calculation has three layers. The first is the revenue not realized during the downtime, obtained by dividing average daily revenue by business hours. The second is the recovery cost: overtime, reprocessing, manual reconciliations, dealing with angry customers. The third, almost always forgotten, is the opportunity cost of the managers who spent two days managing a crisis instead of running the business.

Added together, these layers usually multiply by three or four the initial estimate the board had in mind. And there is a fourth layer, hard to price but easy to recognize: the customer who did not complain, but simply tried the competitor during the downtime and liked the experience. That cost shows up in the following quarter, without an IT incident label.

What evidence does a buyer or investor look for to assess whether the IT operation is institutional or personal?

Technical due diligence looks for verifiable signs, not statements. An up-to-date inventory of systems and licenses, vendor contracts with deadlines and named owners, dated runbooks, an incident log with resolution times, evidence of backup restoration testing, and a history of privileged access. The absence of these artifacts does not mean the operation is bad. It means no one can prove it is good, which in investment vocabulary is the same thing.

The impact is direct and measurable. Undocumented risk turns into a price adjustment, a holdback on part of the payment, or a requirement that specific people stay on after closing. Owners who spent twenty years building value find out too late that part of that value was deposited in an employee's head, and not in the company's coffers.

Where does the internal team's responsibility end and the managed back-office support begin?

The boundary needs to be written before the first incident, not during it. The design that works separates by layer and by time of day: monitoring, first response, and containment sit with the 24x7 back-office support; decisions that affect business processes, prioritization among systems, and the relationship with internal departments sit with the internal team. Each type of incident gets a severity, a primary owner, and an escalation path with a name and a phone number, not a generic department.

A gray area at three in the morning is always a contract failure, never a failure of goodwill. That is why the escalation test before signing is worth more than any sales presentation. And the positive side effect is worth recording: when the middle of the night has a defined owner, the internal professional stops being a hostage to their own knowledge and goes back to taking vacations without bringing the laptop.

How do you turn tacit knowledge into a company asset without paralyzing the operation during the transition?

The transition works in waves, not by big bang. Prioritize the processes with the greatest impact and the least coverage, document each one while it is actually being carried out, and validate the document with the most honest test there is: another person executes it following only the text, with the author watching in silence. Whatever had to be asked is exactly what was missing from the write-up.

In parallel, standardize environments and centralize credentials, because documenting chaos only produces chaos manuals. A realistic schedule covers the critical processes in ninety to one hundred and twenty days, without freezing the operation. The sign that the change has taken hold is not the volume of pages produced, it is how naturally a significant incident is resolved by someone who is not the usual person.

Frequently asked questions

What is the bus factor in IT?

The bus factor is the number of people who would have to become unavailable for a system to stop being operable by the company. A bus factor equal to one means there is only one person capable of carrying out that critical process. The lower the number, the greater the business continuity risk.

Does IT co-management replace the internal team?

No. Co-management adds monitoring, first-level support, and 24x7 response on top of the existing structure, with responsibilities divided by contract. The internal team's role becomes strategic, focused on business processes and continuous improvement, instead of permanent on-call duty and handling alerts outside business hours.

How long does it take to document an IT environment without stopping operations?

A wave-based program typically covers critical processes in ninety to one hundred and twenty days, documenting each procedure while it is being executed in the real-world routine. Validation is performed through independent execution, with another person following only the written text. Operations continue running throughout the entire period.

To find out how many of your operation's critical processes currently depend on a single person, Zamak Technologies conducts a Strategic IT Assessment with no obligation.

When Your Company's IT Lives in One Person's Head
Kleber Leal by Zamak Portal August 10, 2026
Share this post
Tags
Archive