Cloud engineering · DevOps · Data pipeline

Azure Cloud Data Pipeline

A production-style cloud data pipeline built to demonstrate modern Azure, DevOps and infrastructure engineering practices, with secretless runtime authentication and least-privilege access.

Overview

I designed and built a production-style cloud data pipeline to demonstrate modern cloud, DevOps and infrastructure engineering practices using Microsoft Azure.

The application uses Python to retrieve data from a public REST API, transform it and store the resulting output in Azure Blob Storage. It is packaged as a Docker container, stored in Azure Container Registry and executed by an Azure Container Apps Job.

Infrastructure & deployment

The Azure infrastructure is provisioned with Terraform and separated into reusable modules for storage, container registry, Container Apps, container jobs, managed identity and Log Analytics. Terraform state is stored remotely in Azure Storage for reliable infrastructure management.

A GitHub Actions CI/CD pipeline automates deployments. Changes pushed to GitHub build a Docker image, push it to Azure Container Registry with a commit-based version tag, then update the Azure Container Apps Job to run that image.

Security & least-privilege access

The running pipeline holds no secrets. I replaced the storage connection string and the ACR admin credentials, which were long-lived and unscoped, with Azure AD identities. Every role assignment is defined in Terraform and scoped to the individual resource rather than the subscription or resource group.

ActorIdentityRoles & scope
Container Apps JobUser-assigned managed identityAcrPull on the registry, Storage Blob Data Contributor on the storage account
GitHub ActionsService principalAcrPush on the registry, Contributor on the job only, Managed Identity Operator on the job's identity
Local developmentDeveloper's Azure CLI loginStorage Blob Data Contributor, granted explicitly

The ACR admin account is disabled and no storage account key is distributed. The Python application authenticates with DefaultAzureCredential, so the same code uses the managed identity in Azure and my CLI login locally.

I chose a user-assigned identity over a system-assigned one because the job pulls its image when it is created. A system-assigned identity would not exist until the job did, so it could not be granted AcrPull in advance.

Next step: replace the CI service principal's client secret with GitHub OIDC workload identity federation, removing the last stored credential.

Monitoring & troubleshooting

Centralised monitoring and logging are provided through Azure Log Analytics. I use KQL to query container execution logs and troubleshoot deployments.

Key technologies

  • Python
  • Docker
  • Microsoft Azure
  • Terraform
  • GitHub Actions
  • Azure Container Registry
  • Azure Container Apps
  • Azure Blob Storage
  • Azure Log Analytics
  • Managed Identity
  • Azure RBAC
  • KQL
  • Git
  • Linux

Engineering practices demonstrated

  • Infrastructure as Code using modular Terraform
  • Remote Terraform state management
  • Containerisation with Docker
  • Automated CI/CD with GitHub Actions
  • Versioned container deployments using Git commit SHA tags
  • Secretless authentication with managed identity
  • Least-privilege Azure RBAC, scoped per resource and managed in Terraform
  • Centralised logging and monitoring
  • KQL-based troubleshooting
  • Azure CLI and infrastructure troubleshooting
  • Separation of application and infrastructure concerns

Architecture

Public REST APIPythonTransformDockerAzure Container RegistryAzure Container Apps JobAzure Blob Storage
GitHubGitHub ActionsDocker Build → Azure Container Registry → Container Apps