Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Orchestrate tasks with Lakeflow Jobs to read and process a sample dataset. In this quickstart, you:
- Create a new notebook and add code to read a sample dataset of property bookings.
- Save the dataset to Unity Catalog.
- Create a new notebook and add code to read the dataset from Unity Catalog, filter it by year, and display the results.
- Create a new job and configure two tasks using the notebooks.
- Run the job and view the results.
Requirements
If your workspace is Unity Catalog-enabled and Serverless Jobs is enabled, by default, the job runs on Serverless compute. You do not need cluster creation permission to run your job with Serverless compute.
Otherwise, you must have cluster creation permission to create job compute or permissions to all-purpose compute resources.
This quickstart reads from samples.wanderbricks.bookings, which is available in every Unity Catalog-enabled workspace, so there is no source data to set up. To write the table that the second notebook reads, you need permission to create a schema in a catalog (the USE CATALOG and CREATE SCHEMA privileges).
To set these permissions, see your Databricks administrator or Unity Catalog privileges reference.
Create notebooks
The following steps create two notebooks to run in this workflow.
Retrieve and save data
To create a notebook that reads the sample dataset and saves it to Unity Catalog:
Click
New in the sidebar, then click Notebook.
Databricks creates and opens a new, blank notebook in your default folder. The default language is the language you most recently used, and the notebook is automatically attached to the compute resource that you most recently used.(Optional) Rename the notebook Retrieve booking data.
If necessary, change the default language to Python.
Copy the following Python code and paste it into the first cell of the notebook. Before you run it, check that
catalogandschemapoint to a location you can write to. The code creates the schema if it doesn't exist.catalog = "main" schema = "example_output" spark.sql(f"CREATE SCHEMA IF NOT EXISTS {catalog}.{schema}") bookings = spark.read.table("samples.wanderbricks.bookings") bookings.write.mode("overwrite").saveAsTable(f"{catalog}.{schema}.bookings")
Read and display filtered data
To create a notebook that filters and displays your data:
Click
New in the sidebar, then click Notebook.(Optional) Rename the notebook Filter booking data.
The following Python code reads the table you saved in the previous step and creates a temporary view. It also creates a widget that you can use to filter the data in the view by check-in year. Use the same
catalogandschemavalues that you used in the first notebook.from pyspark.sql.functions import year catalog = "main" schema = "example_output" bookings = spark.read.table(f"{catalog}.{schema}.bookings") bookings.createOrReplaceTempView("bookings_table") years = spark.sql("SELECT DISTINCT year(check_in) AS year FROM bookings_table").toPandas()["year"].tolist() years.sort() dbutils.widgets.dropdown("year", "2025", [str(x) for x in years]) display(bookings.filter(year(bookings.check_in) == dbutils.widgets.get("year")))
Create a job
The job you are creating consists of two tasks.
To create the first task:
- In your workspace, click
Jobs & Pipelines in the sidebar.
- Click Create, then Job.
- Click the Notebook tile to configure the first task. If the Notebook tile is not available, click Add another task type and search for Notebook.
- (Optional) Replace the name of the job, which defaults to
New Job <date-time>, with your job name. - In the Task name field, enter a name for the task; for example, retrieve-bookings.
- If necessary, select Notebook from the Type drop-down menu.
- In the Source drop-down menu, select Workspace, which allows you to use a notebook you saved previously.
- For Path, use the file browser to find the first notebook you created, click the notebook name, and click Confirm.
- Click Save task. A notification appears in the upper-right corner of the screen.
To create the second task:
- Click
Add task > Notebook.
- In the Task name field, enter a name for the task; for example, filter-bookings.
- For Path, use the file browser to find the second notebook you created, click the notebook name, and click Confirm.
- Click Add under Parameters. In the Key field, enter
year. In the Value field, enter2025. - Click Save task.
Run the job
To run the job immediately, click
in the upper right corner.
View run details
Click the Runs tab and click the link in the Start time column to open the run that you want to view.
Click either task to see the output and details. For example, click the filter-bookings task to view the output and run details for the filter task:

Run with different parameters
To re-run the job and filter bookings for a different year:
- Click
next to Run now and select Run now with different settings. - In the Value field, enter
2024. - Click Run.
Additional resources
- Configure job settings such as schedules, notifications, and parameters. See Configure and edit Lakeflow Jobs.
- Explore the available task types. See Types of tasks.
- Automatically run jobs on a schedule or when new data arrives. See Automate jobs with schedules and triggers.
- Monitor job runs and set up alerts. See Monitor Lakeflow Jobs.