Bundle packaging modes + ordered multi-bundle deploy - #73
zanitarahimi wants to merge 10 commits into
Conversation
ExecutePipeline emits run_job_task.job_id = ${resources.jobs.X.id},
which only resolves when X is a job in this bundle. In a multi-pipeline
migration each ADF pipeline becomes its own bundle, so a ref to a
sibling pipeline points at a node that does not exist here and
`bundle deploy` fails with "no such node resources.jobs.X".
Add _rewrite_cross_bundle_run_job_refs: before databricks.yml/resource
YAML are written, rewrite run_job_task refs to out-of-bundle jobs into
${var.X} and register X in _cross_bundle_variables (which the existing
YAML builder declares). Operator supplies the numeric job id at deploy
via --var, as SETUP.md documents. Recurses into for_each_task bodies.
This is the stopgap that #10 (ordered cross-pipeline deploy from
control lineage) builds on and keeps as the fallback for unresolved
callees.
Closes #23
Co-authored-by: Isaac
The rewrite count was returned but discarded at both call sites. The function's real output is its in-place mutation of cross_bundle_variables (declared in databricks.yml and surfaced in SETUP.md), so return None and remove the dead counter. Co-authored-by: Isaac
- Remove _rewrite_cross_bundle_run_job_refs from dab_writer: it was
replaced by _rewrite_cross_bundle_job_references during the main
rebase and had no remaining callers. Drop its now-unused
CROSS_BUNDLE_JOB_ID_REF import.
- Fix stale comments that named the removed function and claimed the
old bare ${var.X} scheme (pipeline_graph.py, deployer.py docstring);
the live scheme is ${var.X_job_id}.
- deployer.run(): check for the `databricks` CLI on PATH up front and
return an actionable error instead of an uncaught FileNotFoundError.
Skipped for --dry-run, which never shells out. Add tests.
Declare the union of every workflow's hoisted globals in databricks.yml + SETUP.md, but pass each job only its own workflow's globals. Single-workflow bundles unchanged. Adds TestHoistedGlobalsAcrossGroupedWorkflows.
ghanse
left a comment
There was a problem hiding this comment.
This looks good. I left a few comments. It would be good to add documentation for this feature.
ghanse
left a comment
There was a problem hiding this comment.
Strong, well-tested PR overall. One change I'd like before approving:
deploy_writer.py — single-bundle DEPLOY.md. It tells the operator there's no cross-bundle ordering to worry about and to just run databricks bundle deploy, but a single bundle can still carry a ${var.<callee>_job_id} reference to a pipeline outside the migration (declared with no default). In that case a plain deploy fails on the unset var, and flowx deploy errors with MissingDependencyError unless --allow-missing-deps. Please add a note covering that external-reference case (see the inline comment).
The other inline comments — the duplicate topo sort / CycleError, and the loud-failure test gaps — are non-blocking.
ghanse
left a comment
There was a problem hiding this comment.
Looks very good overall. One finding:
dab_writer.py _namespace_bundle_artifacts breaks dbt-factory pipelines in grouped modes. Step 2 prefixes every notebook path including the resources/*.py PyDABs hook modules, but _collect_pydabs_resource_entries still registers the old resources.<module> path. The file lands at resources/<prefix>/<module>.py while python.resources points at resources.<module>, and bundle deploy can't import the hook.
Confirmed this with a small repro: a dbt pipeline + any second pipeline in a single/per-group bundle is enough because namespacing fires whenever there's more than 1 workflow. The airflow path (_namespace_workflow_assets) skips resources/ for this reason. Needs a fix + a test.
Hey @ghanse on the dbt hook namespacing fix: do you want me to just leave the hook module where it is (so resources/.py still matches python.resources), or namespace it like _namespace_workflow_assets does? |
Namespace it. We may want to reuse |
Bundle packaging modes + ordered multi-bundle deploy
Implements the bundle-packaging design and the ordered auto-deploy follow-up.
Packaging modes
--packaging-mode:per-pipeline(default) /single/per-group; per-group supports--group-by inferred(Run Pipeline call graph) and--group-by spec.bundler/pipeline_graph.py.DEPLOY.mdrecords the suggested callees-first deploy order for every mode.SETUP.md.Ordered deploy
flowx deploy(bundler/deployer.py) discovers the bundles, orders them callees-first, deploys each withdatabricks bundle deploy, reads each deployed job's numeric id frombundle summary, and injects it into callers via--var <callee>_job_id=<id>— no manual job-id wiring.Global-param hoisting across grouped workflows
single/per-group), global-parameter hoisting is union at the bundle level (databricks.ymlvariables:+SETUP.md) and per-workflow at each job (a widget binds to${var.X}only in pipelines that declare X). Single-workflow bundles are byte-identical to before. AddsTestHoistedGlobalsAcrossGroupedWorkflows.Testing
make fmt(ruff + mypy) clean;make testandmake integrationgreen.${var.<callee>_job_id}resolves to the deployed job id.