Is Pulumi a worthy contender in the IaC race? Part 2: from chaos to clarity

avatar
Jeevanandham Poongavanam
11 Sept 2025
  • Share

From debugging nightmares to a “great refactoring,” discover the real lessons we learned scaling Pulumi in production.

This is Part 2 of a three-part series on our production experience with Pulumi. Read Part 1 here for context on why we chose Pulumi and how it compares to alternatives in 2025.

The chaos of learning Pulumi in production

In Part 1, I explained why Pulumi seemed like the perfect choice for our team. Now let me take you back to the initial phase marked by equal parts excitement and frustration. What started as enthusiasm for "infrastructure as real code" quickly descended into a debugging nightmare that had us questioning every decision. But these early struggles would prove invaluable, teaching us lessons that no documentation could provide.

Early mistakes: The single-file monster

Our first mistake was thinking like developers rather than infrastructure engineers. We started with enthusiasm, writing all our infrastructure in a single Python file. What began as a neat 200-line script evolved into a 3,000-line monster that made our IDEs weep. VSCode's IntelliSense would take minutes to load, and finding specific resources became an archaeological expedition.

# How we started - everything in one file
import pulumi
from pulumi_azure_native import resources, storage, network, web
# ... 50 more imports

# 3000+ lines of resources, all in one place
resource_group = resources.ResourceGroup("rg-prod")
vnet = network.VirtualNetwork("vnet-prod", ...)
storage_account = storage.StorageAccount("stgprod", ...)
# ... hundreds more resources
python

The output object confusion

Our biggest challenge, and the source of our most spectacular failure, was understanding Pulumi's Output objects. Coming from imperative programming, we expected to construct values from resource properties like any other variable. This fundamental misunderstanding cost us a lot of time initially.

Here's the exact pattern that caused us grief:

# What we thought would work
storage_account = storage.StorageAccount(
    "storage_account",
    account_name=f"{environment}stgabc",
    resource_group_name=resource_group.name,
    # ... other properties
)

# Using Utility function that creates an output which need to be resolved later
def get_connection_string(resource_group_name: pulumi.Output[str], account_name: pulumi.Output[str]) -> pulumi.Output[str]:
    keys = storage.list_storage_account_keys_output(resource_group_name=resource_group_name, account_name=account_name)
    """
    Connection string contains the account key which is a secret
    """
    # Incorrect usage
    # connection_string = f"DefaultEndpointsProtocol=https;AccountName={account_name};AccountKey={keys.keys[0].value};EndpointSuffix=core.windows.net"

    return pulumi.Output.secret(pulumi.Output.format("DefaultEndpointsProtocol=https;AccountName={0};AccountKey={1};EndpointSuffix=core.windows.net", account_name, keys.keys[0].value))

# If you want to consume the connection string then you have to resolve it and use it from within the lambda function as below
connection_string = get_connection_string() # It is an output object

# Most pulumi resources take the output object as input but when the actual value is needed then we need to explicitly resolve it and consume within the apply method.
connection_string.apply(lambda con_str: ....) # Process the connection string value here
python

The issue? Pulumi's asynchronous execution model means resource properties aren't immediately available — they're wrapped in Output objects that resolve during deployment. Our attempts to use these values directly resulted in connection strings like "DefaultEndpointsProtocol=https;AccountName=<pulumi.output.Output object at 0x...>".

This experience crystallised a crucial insight: while Pulumi lets us use familiar Python syntax, we're not writing a typical Python application — we're declaring a desired state that Pulumi's engine will realise asynchronously. The beauty of using a real programming language can obscure this fundamental paradigm shift.

Once we realised that we're writing a specification for future infrastructure rather than imperatively manipulating resources, everything clicked into place.

The state corruption incident

Our most harrowing experience came six months into production. Our Azure Service Principal credentials expired over a weekend, and our automated pipeline attempted a deployment with invalid credentials. What should have been a simple fix (updating the credentials) turned into a day-long recovery operation.

When we updated the Pulumi configuration with new credentials and ran the pulumi up command, the preview refresh still used the cached expired credentials, resulting in state corruption. The expired cloud provider credentials stored as secrets in the state file were still being used, resulting in authorisation issues. Our attempts to fix this led to a cascade of errors.

The solution was convoluted: We had to downgrade the Azure Native provider, create a new provider instance with fresh credentials, run a careful preview-refresh, then upgrade back to the latest provider version. The inconsistency was maddening — sometimes this process worked flawlessly, other times we had to manually edit the state file.

This incident taught us three critical lessons:

  1. Always monitor credential expiration dates
  2. Understand how Pulumi caches provider configurations
  3. Keep backups of your state files (even when using Pulumi's managed backend)

We have since added a dedicated step in the CI/CD pipeline to verify the credentials before we run the Pulumi CLI to avoid any state corruption. This pipeline design directly addresses the lessons learned from our credential expiration incident by incorporating automated credential monitoring, state backups, and robust error handling at every stage.

Pulumi CI/CD Pipeline

Evolution: From anti-patterns to production-ready

After six months of struggling with our monolithic mess, we reached a breaking point. It had become hard to debug any issues and we were losing control over the code. It was time for a radical transformation — not just of our code, but of our entire approach to infrastructure as code. Here is how we refactored it.

The great refactoring

After six months of pain, we knew something had to change. We embarked on what we called "The great refactoring" — a complete reorganisation of our infrastructure code.

Before: The monolith

infrastructure/
├── __main__.py (3000+ lines)
├── requirements.txt
└── Pulumi.yaml

After: Modular architecture

infrastructure/
├── __main__.py (200 lines - orchestration only)
├── components/
│   ├── networking/
│   │   ├── __init__.py
│   │   ├── NetworkingComponent.py
│   │   └── NetworkingComponentArgs.py
│   ├── storage/
│   │   ├── __init__.py
│   │   └── StorageComponent.py
│   ├── ai_services/
│   └── openfga/
├── custom_providers/
│   └── env_provider/
├── templates/
└── utilities.py

Embracing modularity and order with ComponentResources

The key to our transformation was discovering Pulumi's ComponentResource pattern. Instead of managing hundreds of individual resources, we created logical groupings that encapsulated related infrastructure:

class NetworkingComponent(pulumi.ComponentResource):
    def __init__(self, name: str, args: NetworkingComponentArgs, opts=None):
        super().__init__("pkg:index:NetworkingComponent", name, {}, opts)

        # Create virtual network with all subnets
        self.virtual_network = network.VirtualNetwork(
            f"{args.prefix_name}-vnet-{args.environment}",
            resource_group_name=args.resource_group.name,
            address_space=network.AddressSpaceArgs(
                address_prefixes=[args.network_cidr],
            ),
            opts=pulumi.ResourceOptions(parent=self)
        )

        # Create subnets with proper configurations
        self.web_app_subnet = network.Subnet(
            f"{args.prefix_name}-subnet-webapp",
            virtual_network_name=self.virtual_network.name,
            # Delegation for Azure App Service
            delegations=[network.DelegationArgs(
                name="webapp-delegation",
                service_name="Microsoft.Web/serverFarms"
            )],
            opts=pulumi.ResourceOptions(parent=self)
        )

        # Export the resources for other components
        self.register_outputs({
            "virtual_network": self.virtual_network,
            "web_app_subnet": self.web_app_subnet
        })
python

This approach provided several benefits:

  • Reusability: Components could be instantiated multiple times with different configurations.
  • Encapsulation: Complex logic was hidden behind simple interfaces.
  • Maintainability: Changes to the networking setup required updating only one component.
  • Testing: Components could be unit tested independently.

ComponentResources transformed our thinking from managing individual resources to defining business capabilities. A "NetworkingComponent" became our entire networking capability, not just VNets and subnets. This shift from resource-centric to capability-centric design made our infrastructure mirror our business architecture, turning chaos into a true reflection of our system design.

From output confusion to promise-based clarity

We finally understood how to work with Pulumi's Output objects properly:

from pulumi_azure_native import storage, resources, insights
import pulumi

# Example 1: Using Output.concat for string construction
def get_storage_connection_string(resource_group_name, account_name):
    """Properly construct a connection string using Outputs"""
    keys = storage.list_storage_account_keys_output(
        resource_group_name=resource_group_name,
        account_name=account_name
    )

    # Use Output.concat for string construction with secrets
    return pulumi.Output.secret(
        pulumi.Output.concat(
            "DefaultEndpointsProtocol=https;AccountName=",
            account_name,
            ";AccountKey=",
            keys.keys[0].value,
            ";EndpointSuffix=core.windows.net"
        )
    )

# Example 2: Using Output.all to wait for multiple outputs
def verify_domain:
    ...

comserve_output = pulumi.Output.all(
    resource_group.name, email_comms_service.name, noreply_domain.name
).apply(lambda args: verify_domain(args[0], args[1], args[2]))

# Example 3: Chaining operations with apply
resource_group = resources.ResourceGroup("example-rg")
storage_account = storage.StorageAccount(
    "examplestorage",
    resource_group_name=resource_group.name,
    sku=storage.SkuArgs(name="Standard_LRS")
    ...
)

# Chain operations using apply
storage_endpoint = storage_account.primary_endpoints.apply(
    lambda endpoints: endpoints.blob if endpoints else None
)
python

Key insights about Output handling:

  1. Output.concat vs Output.format: Use concat for simple string joining, format for complex templates.
  2. Output.all: Essential when you need multiple values resolved before proceeding.
  3. Output.secret: Always wrap sensitive values to ensure they're encrypted in state
  4. Apply chaining: Transform outputs through multiple stages while maintaining the async chain.
  5. Conditional outputs: Use Python's native conditionals, but remember the resources inside still create Outputs.

The mental model that finally clicked for us: Outputs are promises for future values. Just like JavaScript promises, you can't access the value directly — you must use apply (conceptually similar to JavaScript’s Promise.then()) to work with the resolved value. Once we embraced this async mindset, Output handling became second nature.

Results of our evolution

While we don't have precise metrics to share, the transformation was dramatic:

  • Deployment times reduced considerably
  • Code reusability increased dramatically — we now share components across five different projects, eliminating thousands of lines of duplicate code
  • Maintainability improved significantly — bug fixes that once required searching through 3,000 lines now involve updating a single component

Most importantly, our development team now actively contributes to infrastructure improvements rather than viewing it as "DevOps territory."

Where we go from here

Our transformation from infrastructure chaos to well-architected design patterns taught us that Pulumi's power comes with responsibility. The flexibility to use real programming languages means you need real software engineering practices. But when you get it right, the results speak for themselves.

This is just the beginning of our Pulumi story. In Part 3, I'll share:

  • The four critical best practices that transformed our Pulumi experience
  • Advanced patterns we've developed for multi-environment deployments
  • Our final verdict: When Pulumi wins and when to consider alternatives
  • Specific recommendations for teams evaluating IaC tools

For now, if you're considering Pulumi or struggling with your current implementation, remember: the chaos is temporary, but the patterns you develop will serve you for years. Start with ComponentResources, respect Outputs, keep your code deterministic and organise from the beginning.

The journey from that first 3,000-line monster file to our current streamlined infrastructure wasn't easy. But today, when a developer says: "I need to add a new service," we don't cringe — we hand them a component template and watch them ship infrastructure as confidently as application code.

That transformation? That's the real power of infrastructure as real code.

But wait - there's more.

Nearform publishes real-world learnings on data & AI, engineering, and digital strategy - with more merged in weekly.

Insights

Perspectives on AI in engineering, product development, and strategy, for enterprise executives.

Community

Deep dives and tutorials by engineers, for engineers.


You may also like

Insight, imagination and expertly engineered solutions to accelerate and sustain progress.