

Building Scalable Serverless Data Pipelines on Google Cloud
Learn how to build scalable serverless data pipelines on Google Cloud using Cloud Run, Pub/Sub, Eventarc, Cloud Storage and BigQuery.

Cem Bakar
Cloud Architect
Cloud Architect focused on building scalable, high-performance systems that power data-driven products and intelligent applications. Bridges business needs with robust, secure cloud architectures.
Modern data architecture needs to scale without creating unnecessary infrastructure management. Serverless data pipelines help organizations process, transform and move data without provisioning or maintaining servers.
A serverless pipeline is the complete event-driven workflow. Individual Google Cloud services provide the compute, messaging, routing and storage capabilities that execute each stage of that workflow. Understanding how these services work together is critical for cloud architects and data teams building reliable, scalable systems.
What Is a Serverless Data Pipeline?
A serverless data pipeline is a workflow that moves data from a source to a destination using managed cloud services. The cloud provider manages the underlying infrastructure, while the development team defines how data is collected, processed, routed and stored.
Serverless does not mean that no servers are involved. It means your team does not need to provision or manage those servers directly. Services can run on demand and scale based on workload, configuration and platform limits.
Connecting a Serverless Pipeline to Google Cloud Services
Designing a serverless pipeline starts with mapping each logical stage to the appropriate Google Cloud service.
The ingestion and routing stage captures events and moves them to the next part of the workflow. Pub/Sub can provide reliable, asynchronous messaging between systems, while Eventarc can route events from Google Cloud services to destinations such as Cloud Run.
The compute and processing stage handles the business logic. Cloud Run functions, formerly called Cloud Functions, are well suited to lightweight, single-purpose code triggered by an event. Cloud Run services are better suited to containerized workloads that require custom dependencies, more processing control or concurrent request handling.
The storage and analytics stage provides a destination for raw and processed data. Cloud Storage can hold files and unstructured data, while BigQuery supports structured data analysis at scale.
Together, these services make it possible to build a complete serverless data pipeline without managing the underlying servers.
Real-World Pipeline: API to Data Warehouse Ingestion
Consider a practical data engineering scenario where you need to ingest daily project management data from a third-party REST API and load it into a data warehouse. The pipeline begins with Cloud Scheduler, which triggers the workflow on a daily basis.
This schedule invokes a Python-based Cloud Function. The function authenticates with the external API, retrieves the JSON payload and performs initial schema validation. To ensure data durability, the function writes this raw JSON directly to a Cloud Storage bucket, creating an immutable audit trail.
The creation of this file in Cloud Storage acts as an event trigger, routing a message to a Cloud Run service. The Cloud Run container takes over the heavy lifting: it parses the raw JSON, transforms the nested fields into a flattened structure and streams the cleaned data directly into BigQuery tables. The entire flow operates without provisioning a single server, scaling exactly to the size of the daily payload.
Example: Building an API-to-BigQuery Data Pipeline
Consider a data engineering scenario where a team needs to retrieve daily project management data from a third-party REST API and load it into a data warehouse.
The pipeline begins with Cloud Scheduler, which starts the workflow at a set time each day.
The scheduled request invokes a Python-based Cloud Run function. The function authenticates with the external API, retrieves the JSON payload and performs an initial schema validation.
The function then writes the original JSON to a Cloud Storage bucket. Preserving the raw data creates an audit record that can support troubleshooting, reprocessing and historical analysis. Additional controls, such as retention policies or object versioning, can be configured when stronger protection is required.
When the file is created in Cloud Storage, Eventarc routes the event to a Cloud Run service. The Cloud Run container handles the heavier processing tasks. It parses the JSON, transforms nested fields into a flattened structure and writes the cleaned data to BigQuery tables.
The workflow runs on demand without requiring a server to remain active between scheduled jobs. Each component can also be monitored, secured and scaled independently.
The completed pipeline looks like this:
Cloud Scheduler → Cloud Run function → Cloud Storage → Eventarc → Cloud Run service → BigQuery
For another practical example, see how Napkyn approaches an API-to-BigQuery data pipeline using YouTube data.
Common Uses for Google Cloud Serverless Services
Each component of the pipeline can also operate independently within broader cloud architectures.
Cloud Run Functions for Event-Driven Tasks
Cloud Run functions are useful for lightweight, single-purpose tasks triggered by an event.
For example, when an application uploads a raw image to a Cloud Storage bucket, a function can automatically run a Python script that crops and compresses the image. It can then save the optimized version to a separate bucket for use by a website or application.
This approach works well when the task has a clear trigger, performs a focused action and does not require a complex application environment.
Cloud Run for Containerized Applications
Cloud Run is a strong option when a workload requires containerization, custom dependencies or concurrent request handling.
A common use case is hosting a backend API. Cloud Run can create additional instances as incoming traffic increases and scale down when demand falls. Teams can also configure minimum and maximum instance limits based on performance and cost requirements.
Cloud Run can support APIs, data transformation services, scheduled jobs and other containerized workloads without requiring teams to manage a traditional server environment.
Cloud Storage as a Data Landing Zone
Cloud Storage does more than hold files. It often acts as the landing zone for serverless data pipelines.
For example, an automated system might place daily CSV exports of website or customer data into a specific bucket. The bucket then becomes the reliable starting point for downstream extraction, transformation and loading processes.
Keeping raw data in Cloud Storage also allows teams to reprocess files if transformation logic changes or if an error occurs later in the pipeline.
Pub/Sub for Asynchronous Messaging
Pub/Sub helps decouple applications and services through asynchronous messaging.
In an e-commerce order system, for example, the website can publish an order event to a Pub/Sub topic. Independent services responsible for inventory, billing and shipping notifications can subscribe to that topic and process the event separately.
The website does not need to wait for every backend process to finish before confirming the order. This helps keep the checkout experience responsive while allowing downstream systems to operate independently.
Why Use Serverless Data Pipelines on Google Cloud?
Serverless data pipelines can help teams:
Reduce infrastructure management
Scale individual components based on demand
Process events and data asynchronously
Separate ingestion, transformation and storage responsibilities
Pay for resources when workloads are running
Build modular systems that are easier to update
Create reliable paths from operational data to analytics platforms such as BigQuery
The right architecture depends on the workload. Lightweight event processing may only require a Cloud Run function, while larger data workflows may combine Cloud Scheduler, Eventarc, Pub/Sub, Cloud Run, Cloud Storage and BigQuery.
The key is to start with the workflow rather than the individual products. Once the stages of the pipeline are clear, each Google Cloud service can be selected based on the role it needs to perform.
Need Help Building a Serverless Data Pipeline?
Napkyn helps organizations design and build scalable data pipelines using Google Cloud, BigQuery and other marketing and business data sources. Whether you are replacing manual processes, connecting disconnected systems or creating a stronger foundation for analytics and AI, we can help you identify the right architecture and put it into practice.
Contact us to discuss your Google Cloud and data architecture needs.
Napkyn is a Google Cloud Partner and Google Marketing Platform Sales Partner.
More Insights


Building Scalable Serverless Data Pipelines on Google Cloud

Cem Bakar
Cloud Architect
Oct 7, 2026
Read More


Understanding the IAB TCF v2.3 Update: New Vendor Disclosure Requirements Explained

Rob English
Director, Data Solutions
Jun 17, 2026
Read More


CM360 Path to Conversion Reporting: Full Customer Journey Attribution

Monika Boldak
Associate Director, Marketing
Jun 3, 2026
Read More
More Insights
Sign Up For Our Newsletter

Napkyn Inc.
204-78 George Street, Ottawa, Ontario, K1N 5W1, Canada
Napkyn US
6 East 32nd Street, 9th Floor, New York, NY 10016, USA
212-247-0800 | info@napkyn.com

Building Scalable Serverless Data Pipelines on Google Cloud
Learn how to build scalable serverless data pipelines on Google Cloud using Cloud Run, Pub/Sub, Eventarc, Cloud Storage and BigQuery.

Cem Bakar
Cloud Architect
October 7, 2026
Cloud Architect focused on building scalable, high-performance systems that power data-driven products and intelligent applications. Bridges business needs with robust, secure cloud architectures.
Modern data architecture needs to scale without creating unnecessary infrastructure management. Serverless data pipelines help organizations process, transform and move data without provisioning or maintaining servers.
A serverless pipeline is the complete event-driven workflow. Individual Google Cloud services provide the compute, messaging, routing and storage capabilities that execute each stage of that workflow. Understanding how these services work together is critical for cloud architects and data teams building reliable, scalable systems.
What Is a Serverless Data Pipeline?
A serverless data pipeline is a workflow that moves data from a source to a destination using managed cloud services. The cloud provider manages the underlying infrastructure, while the development team defines how data is collected, processed, routed and stored.
Serverless does not mean that no servers are involved. It means your team does not need to provision or manage those servers directly. Services can run on demand and scale based on workload, configuration and platform limits.
Connecting a Serverless Pipeline to Google Cloud Services
Designing a serverless pipeline starts with mapping each logical stage to the appropriate Google Cloud service.
The ingestion and routing stage captures events and moves them to the next part of the workflow. Pub/Sub can provide reliable, asynchronous messaging between systems, while Eventarc can route events from Google Cloud services to destinations such as Cloud Run.
The compute and processing stage handles the business logic. Cloud Run functions, formerly called Cloud Functions, are well suited to lightweight, single-purpose code triggered by an event. Cloud Run services are better suited to containerized workloads that require custom dependencies, more processing control or concurrent request handling.
The storage and analytics stage provides a destination for raw and processed data. Cloud Storage can hold files and unstructured data, while BigQuery supports structured data analysis at scale.
Together, these services make it possible to build a complete serverless data pipeline without managing the underlying servers.
Real-World Pipeline: API to Data Warehouse Ingestion
Consider a practical data engineering scenario where you need to ingest daily project management data from a third-party REST API and load it into a data warehouse. The pipeline begins with Cloud Scheduler, which triggers the workflow on a daily basis.
This schedule invokes a Python-based Cloud Function. The function authenticates with the external API, retrieves the JSON payload and performs initial schema validation. To ensure data durability, the function writes this raw JSON directly to a Cloud Storage bucket, creating an immutable audit trail.
The creation of this file in Cloud Storage acts as an event trigger, routing a message to a Cloud Run service. The Cloud Run container takes over the heavy lifting: it parses the raw JSON, transforms the nested fields into a flattened structure and streams the cleaned data directly into BigQuery tables. The entire flow operates without provisioning a single server, scaling exactly to the size of the daily payload.
Example: Building an API-to-BigQuery Data Pipeline
Consider a data engineering scenario where a team needs to retrieve daily project management data from a third-party REST API and load it into a data warehouse.
The pipeline begins with Cloud Scheduler, which starts the workflow at a set time each day.
The scheduled request invokes a Python-based Cloud Run function. The function authenticates with the external API, retrieves the JSON payload and performs an initial schema validation.
The function then writes the original JSON to a Cloud Storage bucket. Preserving the raw data creates an audit record that can support troubleshooting, reprocessing and historical analysis. Additional controls, such as retention policies or object versioning, can be configured when stronger protection is required.
When the file is created in Cloud Storage, Eventarc routes the event to a Cloud Run service. The Cloud Run container handles the heavier processing tasks. It parses the JSON, transforms nested fields into a flattened structure and writes the cleaned data to BigQuery tables.
The workflow runs on demand without requiring a server to remain active between scheduled jobs. Each component can also be monitored, secured and scaled independently.
The completed pipeline looks like this:
Cloud Scheduler → Cloud Run function → Cloud Storage → Eventarc → Cloud Run service → BigQuery
For another practical example, see how Napkyn approaches an API-to-BigQuery data pipeline using YouTube data.
Common Uses for Google Cloud Serverless Services
Each component of the pipeline can also operate independently within broader cloud architectures.
Cloud Run Functions for Event-Driven Tasks
Cloud Run functions are useful for lightweight, single-purpose tasks triggered by an event.
For example, when an application uploads a raw image to a Cloud Storage bucket, a function can automatically run a Python script that crops and compresses the image. It can then save the optimized version to a separate bucket for use by a website or application.
This approach works well when the task has a clear trigger, performs a focused action and does not require a complex application environment.
Cloud Run for Containerized Applications
Cloud Run is a strong option when a workload requires containerization, custom dependencies or concurrent request handling.
A common use case is hosting a backend API. Cloud Run can create additional instances as incoming traffic increases and scale down when demand falls. Teams can also configure minimum and maximum instance limits based on performance and cost requirements.
Cloud Run can support APIs, data transformation services, scheduled jobs and other containerized workloads without requiring teams to manage a traditional server environment.
Cloud Storage as a Data Landing Zone
Cloud Storage does more than hold files. It often acts as the landing zone for serverless data pipelines.
For example, an automated system might place daily CSV exports of website or customer data into a specific bucket. The bucket then becomes the reliable starting point for downstream extraction, transformation and loading processes.
Keeping raw data in Cloud Storage also allows teams to reprocess files if transformation logic changes or if an error occurs later in the pipeline.
Pub/Sub for Asynchronous Messaging
Pub/Sub helps decouple applications and services through asynchronous messaging.
In an e-commerce order system, for example, the website can publish an order event to a Pub/Sub topic. Independent services responsible for inventory, billing and shipping notifications can subscribe to that topic and process the event separately.
The website does not need to wait for every backend process to finish before confirming the order. This helps keep the checkout experience responsive while allowing downstream systems to operate independently.
Why Use Serverless Data Pipelines on Google Cloud?
Serverless data pipelines can help teams:
Reduce infrastructure management
Scale individual components based on demand
Process events and data asynchronously
Separate ingestion, transformation and storage responsibilities
Pay for resources when workloads are running
Build modular systems that are easier to update
Create reliable paths from operational data to analytics platforms such as BigQuery
The right architecture depends on the workload. Lightweight event processing may only require a Cloud Run function, while larger data workflows may combine Cloud Scheduler, Eventarc, Pub/Sub, Cloud Run, Cloud Storage and BigQuery.
The key is to start with the workflow rather than the individual products. Once the stages of the pipeline are clear, each Google Cloud service can be selected based on the role it needs to perform.
Need Help Building a Serverless Data Pipeline?
Napkyn helps organizations design and build scalable data pipelines using Google Cloud, BigQuery and other marketing and business data sources. Whether you are replacing manual processes, connecting disconnected systems or creating a stronger foundation for analytics and AI, we can help you identify the right architecture and put it into practice.
Contact us to discuss your Google Cloud and data architecture needs.
Napkyn is a Google Cloud Partner and Google Marketing Platform Sales Partner.
More Insights

Building Scalable Serverless Data Pipelines on Google Cloud

Cem Bakar
Cloud Architect
Oct 7, 2026
Read More

The Most Durable AI Investment: A Strong Data Foundation

Jasmine Libert
Senior VP Growth and GTM
Sep 28, 2026
Read More

How to Detect Google Analytics Tracking Issues Early with Automated QA

Skylar van Dalen-Flude
Senior Data Analyst
Aug 26, 2026
Read More
More Insights
Sign Up For Our Newsletter



