Cloud Infrastructure & DevOps: the questions leadership should ask before investing

The board-level questions that turn Cloud Infrastructure & DevOps from a technical purchase into a controlled business decision.

KANJ Advisory Team 7 min read
Cloud Infrastructure & DevOps: the questions leadership should ask before investing

The questions that help leadership understand whether cloud infrastructure is delivering the resilience, control, flexibility and value the business actually needs.

Cloud infrastructure decisions have a habit of becoming technology conversations. Platforms, architecture, automation, security and cost all matter, but some of the most important questions come before them.

What is the business trying to improve? What cannot be allowed to fail? What level of cost and dependency is acceptable? And how will leadership know whether the technology is actually delivering what the organisation needs?

These questions matter whether a business is considering a significant new investment or reviewing infrastructure that has evolved over several years.

Leadership does not need to design the architecture. It does, however, need to make sure the people designing and managing it understand what the business is trying to achieve and which compromises it is prepared to make.

There is rarely a single "best" cloud architecture. Greater resilience may require greater investment. More control can create more operational responsibility. Cloud-native services can simplify management while increasing dependency on a particular provider. Security decisions can affect usability and agility.

Cloud infrastructure is therefore partly a technical challenge, but it is also a series of business trade-offs expressed through technology.

1. What constraint are we actually trying to remove?

"Moving to the cloud" is not a business objective. Neither is "implementing DevOps".

The useful question is what the organisation expects to become materially better as a result.

Perhaps the existing environment is slowing product development, making acquisitions difficult to integrate or creating unacceptable recovery risk. Infrastructure costs may be increasing without a corresponding improvement in capability. Expansion into new locations may have exposed limitations that were less important when the business was smaller. Or maintaining existing systems may simply be consuming too much time and money.

The warning signs are not always found in the infrastructure itself. They can appear elsewhere in the business: projects take longer, reporting becomes harder, systems do not integrate easily, teams create manual workarounds, or seemingly straightforward changes require disproportionate effort.

Starting with the constraint rather than the technology changes the conversation. Instead of asking what cloud platform, service or architecture the organisation should buy, leadership can ask what capability the business needs and work backwards from there.

It also creates a useful test for proposed investment. If nobody can clearly explain what should become better, faster, safer or less expensive as a result, it is worth questioning whether the technology is solving the right problem.

2. Do we properly understand what we already have?

Technology environments rarely become unsuitable overnight. Complexity tends to accumulate.

Applications become dependent on ageing components. Cloud resources are created for projects and remain long after those projects have finished. Different systems perform similar functions. Deployments require manual intervention. Monitoring becomes fragmented. Suppliers change, people leave and knowledge of important systems gradually becomes concentrated in a small number of individuals.

None of these necessarily justifies a major infrastructure change on its own.

Together, however, they can make the environment more expensive to operate, harder to secure and increasingly difficult to change.

Before designing something new, it is worth understanding the existing environment properly: applications, integrations, infrastructure, licences, suppliers, security controls, data, operational processes and the dependencies between them.

This is particularly important when infrastructure has grown organically. A business can accumulate considerable technical debt without experiencing a single obvious failure.

Without that understanding, a cloud project risks moving existing complexity into a newer platform rather than removing it.

3. What level of resilience does the business actually need?

"High availability", "redundancy" and "disaster recovery" describe technical capabilities. They do not tell leadership what happens to the business when something fails.

That conversation needs to be more specific.

How long could a customer-facing application be unavailable before there was a serious commercial consequence? How much data could be lost without materially affecting operations? Which systems would need to recover first? Which processes could continue manually? What would happen if a major technology or connectivity provider became unavailable?

The answers will differ across the organisation.

A customer transaction platform may justify significantly greater resilience than an internal application that can be unavailable for several hours without serious consequence. Equally, a relatively unimportant-looking system may support a process that prevents the rest of the business from operating when it fails.

Understanding those dependencies matters because resilience has a cost. Designing everything for maximum availability can be unnecessarily expensive. Designing primarily around cost can leave the organisation accepting risks it never consciously agreed to.

The objective is not simply to eliminate downtime. It is to understand where disruption matters, what the organisation can tolerate and where investment in resilience is justified.

4. Who is responsible when something goes wrong?

Moving infrastructure to a major cloud provider does not transfer responsibility for everything running on it.

Cloud operates on a shared responsibility model. The provider takes responsibility for defined elements of the underlying service, while the customer retains responsibility for areas that can include identities, access, configuration, data and how the services are used.

For many businesses, the practical picture is more complicated still.

There may be an internal IT team, cloud provider, managed service provider, software vendor, development company and security partner involved in different parts of the same environment. Each relationship may work perfectly well until an incident crosses the boundary between them.

Leadership should therefore be able to ask a deceptively simple question:

If this system is compromised, unavailable or incorrectly configured tomorrow, who is responsible for what?

The answer should be clear before an incident occurs.

Good architecture is not simply about where technology sits. Responsibility, ownership, escalation and decision-making need to be understood alongside it.

5. What will this really cost to operate?

Cloud has changed the economics of infrastructure, but it has not made infrastructure inherently inexpensive.

The initial project or migration cost is only one part of the calculation. There may be ongoing expenditure on compute, storage, network traffic, licences, monitoring, security, backup, support, development tooling and the people required to manage the environment.

More importantly, those costs change as the business changes.

An architecture that looks inexpensive at today's workload may behave very differently at scale. Poorly governed environments can accumulate unused resources and duplicated services. Conversely, spending more on automation, managed services or better architecture may reduce the engineering effort required to maintain the environment.

That makes simply asking whether the cloud bill can be reduced a fairly limited question.

A more useful one is:

Are we getting appropriate business value for what we are spending?

Cost optimisation is not necessarily about finding the cheapest possible infrastructure. It is about understanding what the organisation is paying for, whether it still needs it and whether that expenditure supports the level of performance, resilience and flexibility the business requires.

6. Will this make the business easier to change - or harder to leave?

This is where DevOps becomes particularly relevant to leadership.

The board does not need to understand the mechanics of deployment pipelines, infrastructure as code or container orchestration. It should care about what those capabilities allow the organisation to do.

Can software changes be introduced more safely? Can environments be created without weeks of manual work? Can infrastructure changes be made consistently rather than relying on an engineer remembering a sequence of steps? Can problems be identified quickly and failed changes reversed without creating a major incident?

But flexibility has another dimension that is easily overlooked: dependency.

What technologies and suppliers is the organisation becoming dependent upon? Could another competent team operate the environment? Is it properly documented? Is important knowledge concentrated with one supplier or individual? If a commercial relationship changed, could the business realistically regain control or move elsewhere?

Some dependency is inevitable. The objective is not to avoid it entirely, but to understand it.

A useful test of any infrastructure decision is therefore:

If the business needs something different in two years, how difficult will this decision be to change?

Technology should give the organisation options. An environment that works extremely well today but makes tomorrow's change prohibitively difficult may simply be creating a different form of technical debt.

7. How will we know whether the investment worked?

Technology projects are often good at defining what will be delivered and less precise about what should improve as a result.

Those measures are most useful when they are agreed before significant investment is made.

If resilience is the objective, measure availability and recovery performance. If cost is the problem, understand the cost of operating the service and how that changes with demand. If development speed matters, measure how quickly and reliably changes reach production. If operational efficiency is important, consider how much engineering time is spent maintaining infrastructure and resolving incidents.

Not everything needs to become a KPI. But leadership should have enough evidence to determine whether the original problem has actually improved.

A technically successful migration is not necessarily a successful business investment.

An organisation can move every workload, complete the project on time and still struggle to answer the most important question:

What became better because we did it?

The technology should follow the business decision

Cloud architecture involves balancing security, reliability, operational effectiveness, performance, cost and flexibility. Improving one can affect another, which is why infrastructure decisions rarely have a universally correct answer.

The role of the technical team is to understand those requirements and translate them into appropriate architecture.

Leadership has a different responsibility. It needs to understand the consequences of the choices being made.

Before approving a significant Cloud Infrastructure & DevOps investment, or deciding that the current environment remains appropriate, four things should therefore be reasonably clear:

  • What business outcome or constraint are we addressing?

  • What risks, costs and dependencies are we accepting?

  • Who owns the important decisions and operational responsibilities?

  • What evidence will tell us whether the technology is working?

The most useful cloud conversations therefore begin before anyone discusses platforms or architecture. They begin with what the business needs to protect, improve or change - and what it is prepared to spend, own and depend upon to achieve it.

Once those decisions are understood, technology has a much clearer job to do.

Practical checks

  • Name the owners for key decisions and service changes.
  • Agree the evidence leadership needs to review progress.
  • Separate immediate operational fixes from strategic improvements.
  • Set a review rhythm before the work moves into delivery.
Keep exploring

Related insights

let's collaborate

Need IT That Reduces Risk and Stands Up to Regulation?

Let's strengthen reliability and optimise your IT for efficiency.