Cloud engineering · DevOps · Data pipeline
Azure Cloud Data Pipeline
A production-style cloud data pipeline built to demonstrate modern Azure, DevOps and infrastructure engineering practices, with secretless runtime authentication and least-privilege access.
Overview
I designed and built a production-style cloud data pipeline to demonstrate modern cloud, DevOps and infrastructure engineering practices using Microsoft Azure.
The application uses Python to retrieve data from a public REST API, transform it and store the resulting output in Azure Blob Storage. It is packaged as a Docker container, stored in Azure Container Registry and executed by an Azure Container Apps Job.
Infrastructure & deployment
The Azure infrastructure is provisioned with Terraform and separated into reusable modules for storage, container registry, Container Apps, container jobs, managed identity and Log Analytics. Terraform state is stored remotely in Azure Storage for reliable infrastructure management.
A GitHub Actions CI/CD pipeline automates deployments. Changes pushed to GitHub build a Docker image, push it to Azure Container Registry with a commit-based version tag, then update the Azure Container Apps Job to run that image.
Security & least-privilege access
The running pipeline holds no secrets. I replaced the storage connection string and the ACR admin credentials, which were long-lived and unscoped, with Azure AD identities. Every role assignment is defined in Terraform and scoped to the individual resource rather than the subscription or resource group.
| Actor | Identity | Roles & scope |
|---|---|---|
| Container Apps Job | User-assigned managed identity | AcrPull on the registry, Storage Blob Data Contributor on the storage account |
| GitHub Actions | Service principal | AcrPush on the registry, Contributor on the job only, Managed Identity Operator on the job's identity |
| Local development | Developer's Azure CLI login | Storage Blob Data Contributor, granted explicitly |
The ACR admin account is disabled and no storage account key is distributed. The Python application authenticates with DefaultAzureCredential, so the same code uses the managed identity in Azure and my CLI login locally.
I chose a user-assigned identity over a system-assigned one because the job pulls its image when it is created. A system-assigned identity would not exist until the job did, so it could not be granted AcrPull in advance.
Next step: replace the CI service principal's client secret with GitHub OIDC workload identity federation, removing the last stored credential.
Monitoring & troubleshooting
Centralised monitoring and logging are provided through Azure Log Analytics. I use KQL to query container execution logs and troubleshoot deployments.
Key technologies
Engineering practices demonstrated
- Infrastructure as Code using modular Terraform
- Remote Terraform state management
- Containerisation with Docker
- Automated CI/CD with GitHub Actions
- Versioned container deployments using Git commit SHA tags
- Secretless authentication with managed identity
- Least-privilege Azure RBAC, scoped per resource and managed in Terraform
- Centralised logging and monitoring
- KQL-based troubleshooting
- Azure CLI and infrastructure troubleshooting
- Separation of application and infrastructure concerns