Skip to main content
FlareDB runs Apache Beam pipelines as jobs. Build your pipeline with the Beam Python SDK as usual, then set FlareRunner as the runner to submit it to a running FlareDB instance.

Prerequisites

  • A running FlareDB instance. See the Quickstart.
  • Python 3.9+.
  • An Apache Beam pipeline. The examples in this guide use Apache Beam 2.76.0.

1. Install the runner dependency

Create and activate a virtual environment, then install the flaredb-runner package from the repository. It includes flare and apache-beam dependency.

2. Configure the pipeline

Set FlareRunner as the runner and point it at the FlareDB instance:
wordcount.py
Unlike Java, there is nothing to stage manually: FlareDB ships your pipeline code to the workers automatically. See the full pipeline in the WordCount example.

3. Run the pipeline

With a FlareDB instance running, submit the pipeline:

Pipeline options

Pass options as keyword arguments to PipelineOptions: These options can also be provided as command-line arguments:

Pipeline logs

A JOB-ID is generated automatically for each submitted job. Use it to view logs for a specific job, or stream the most recent logs when you do not provide an ID. Usage:
The JOB-ID is logged during job submission via the runner SDK.