Update dependency io.openlineage:openlineage-java to v1.53.0 - #3002
Open
renovate[bot] wants to merge 1 commit into
Open
Update dependency io.openlineage:openlineage-java to v1.53.0#3002renovate[bot] wants to merge 1 commit into
renovate[bot] wants to merge 1 commit into
Conversation
❌ Deploy Preview for peppy-sprite-186812 failed.
|
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
January 20, 2025 19:42
d415db2 to
a641137
Compare
Codecov ReportAll modified and coverable lines are covered by tests ✅
Additional details and impacted files@@ Coverage Diff @@
## main #3002 +/- ##
=========================================
Coverage 81.18% 81.18%
Complexity 1506 1506
=========================================
Files 268 268
Lines 7356 7356
Branches 325 325
=========================================
Hits 5972 5972
Misses 1226 1226
Partials 158 158 ☔ View full report in Codecov by Sentry. 🚀 New features to boost your workflow:
|
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
7 times, most recently
from
January 22, 2025 13:56
403306b to
96a3b7b
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
3 times, most recently
from
February 7, 2025 00:42
dc448cf to
ebaf22c
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
February 25, 2025 18:55
ebaf22c to
a164571
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
March 17, 2025 12:11
a164571 to
f11be42
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
6 times, most recently
from
March 26, 2025 17:23
432f8d4 to
9d7550b
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
3 times, most recently
from
March 27, 2025 07:34
43064c3 to
7c19900
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
April 10, 2025 15:36
7c19900 to
828aa83
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
August 11, 2025 22:00
d45585c to
d31b116
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
October 1, 2025 22:10
d31b116 to
bec3771
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
October 7, 2025 17:45
bec3771 to
cfc84da
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
November 14, 2025 02:13
cfc84da to
f9f7c18
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
November 14, 2025 12:55
f9f7c18 to
3b1b2ea
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
December 11, 2025 13:36
3b1b2ea to
00ac754
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
January 7, 2026 14:39
00ac754 to
e8d4cc6
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
January 8, 2026 23:36
e8d4cc6 to
2d1359e
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
January 23, 2026 01:38
2d1359e to
cd81291
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
February 17, 2026 14:48
cd81291 to
3f63859
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
February 20, 2026 18:12
3f63859 to
96d5e69
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
March 11, 2026 17:27
96d5e69 to
b5a9a65
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
3 times, most recently
from
March 17, 2026 00:14
8c49891 to
5c5c351
Compare
renovate
Bot
force-pushed
the
renovate/openlineageversion
branch
from
April 6, 2026 23:43
5c5c351 to
1f0a6f6
Compare
Signed-off-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR contains the following updates:
1.23.0→1.53.0Release Notes
OpenLineage/OpenLineage (io.openlineage:openlineage-java)
v1.53.0Compare Source
Added
#4759@tnazarewExposes GCP Lineage API retry settings through transport configuration and applies them to the underlying client.
#4798@tnazarewExposes the lineage producer's retry policy through
GcsTransportconfiguration instead of always using defaults.#4778@mobuchowskiGenerates typed Jackson union interfaces and concrete variants, enabling correct construction and deserialization of explicit lineage entries.
#4812@karthikchundi-commitsParses Oracle RAC and failover TNS descriptors into stable
oracle://host:portdataset namespaces.#4747@mobuchowskiAdds Python 3.14 support and validates the Spark integration on Spark 4.2 with Java 17 and Java 21.
#4465@kchledowskiAdds opt-in normalization of input and output datasets before the Python client emits an event.
#4875@mobuchowskiAdds canonical AWS Glue table identifiers as symlinks for datasets emitted by the dbt Athena integration.
#4861@fmorillo7694Aligns SQL and DataStream Kinesis dataset identities and converts Kinesis type metadata into schema facets.
#4878@MSDehghanAdds dataset identifiers plus catalog and storage facets for reads and writes through the official ClickHouse V2 catalog.
#4782@mobuchowski with @tnazarewCaptures expression descriptions while resolving Spark column-level lineage transformation chains.
#4797@tnazarewAdds dedicated handling for Google Lakehouse Catalog identifiers instead of relying on the generic REST catalog path.
#4848@mishrasangeeta87Emits the filesystem path read by
LOAD DATA INPATHas an input dataset alongside the target table output.#4804@mobuchowskiAdds Job and Dataset facets for declaring exact dataset-, field-, and job-level relationships without Cartesian-product inference.
#4859@mattfaltynExtracts table lineage from nested join expressions while preserving aliases, constraints, and subquery traversal.
Changed
#4902@dolfinusMoves
AsyncHttpTransportto the maintainedhttpx2drop-in replacement and updates related configuration and tests.#4819@mobuchowskiAdds Java, Python, and Go validation for schema fields that must be supplied together, including explicit lineage job identities.
Fixed
#4612@matveeysvParses custom ports after bracketed IPv6 hosts without appending an incorrect default port.
#4838@MSDehghanPrevents partition-aware dataset reduction from failing when generated or mocked datasets have null facets.
#4824@mattfaltynKeeps
AsyncHttpTransportusable afterwait_for_completion()so later events are still delivered.#4882@mattfaltynMoves released completion events to a worker-owned backlog so bursts cannot block the sole worker on its bounded queue.
#4900@mattfaltynPrevents terminal run events from being stranded when their
STARTevent completes concurrently.#4509@hcthakur2004Avoids locale-dependent decoding failures by reading Python client YAML configuration explicitly as UTF-8.
#4808@chuenchen309Propagates standardized parent and root run metadata when the dbt wrapper consumes local artifacts.
#4809@chuenchen309Adds the dbt test name to structured-log
dataQualityAssertions, matching the run-results processor.#4777@kacpermudaRestores short assertion types and column associations for generic tests in dbt manifest v12.
#4846@chuenchen309Replaces the unresolved producer placeholder and emits RFC 3339 timestamps with a UTC offset.
#4733@chuenchen309Uses the renamed expectation type field and maps file-size results to
bytes, restoring metrics and assertions.#4811@chuenchen309Allows
OpenLineageValidationActionto construct on Great Expectations 1.x while retaining compatibility with versions that require the argument.#4765@sulikismaylovvUpdates bundled Jackson dependencies across Java, Spark, Flink, and Hive and adjusts shaded jars for the newer release.
#4853@Poojitha-R-RaoApplies the subsequent Jackson patch release consistently across the Java client and integrations.
#4726@zerafachrisApplies configured dataset path removal to RDD job inputs and outputs, matching SQL and DataFrame behavior.
#4894@mobuchowskiEvicts completed jobs, stages, and metrics so long-lived Spark drivers do not retain state without bound.
#4884@MSDehghanRestores output datasets for Spark 4 structured-streaming writes through Delta's V1 sink.
#4850@mishrasangeeta87Emits source inputs and Unity Catalog target outputs for proprietary Databricks
COPY INTOplan variants.#4849@mishrasangeeta87Restores output datasets for CTAS and related V2 create or replace commands on Databricks runtimes.
#4815@mishrasangeeta87Emits the DELETE target as an output and tables referenced by predicate subqueries as inputs.
#4835@mishrasangeeta87Emits UPDATE targets and SET or WHERE subquery inputs from proprietary Databricks plan nodes.
#4898@MSDehghanStops read-only V2 scans, such as lazy Iceberg checkpoints, from being reported as writes.
#4772@MSDehghanRestores output lineage for Spark 4 structured-streaming jobs that use V1 sinks.
#4896@JDarDagranResolves maintenance-action datasets from the underlying Iceberg relation when writes use
SparkCachedTableCatalog.#4779@mobuchowskiAvoids listener failures after Spark 4.2 changed
CatalogManagerfrom a class to an interface.#4775@mattfaltynClassifies explicit DELETE targets as outputs and joined lookup tables as inputs.
#4791@mattfaltynReports source tables as inputs and the destination location as output for Snowflake unload statements.
#4865@mattfaltynCaptures predicate columns and subquery tables referenced only by aggregate
FILTERexpressions.#4763@mattfaltynIncludes tables referenced by BigQuery
ARRAYsubqueries in input lineage.#4785@mattfaltynIncludes tables read only by subqueries in
HAVINGexpressions.#4783@mattfaltynIncludes tables read only by subqueries in
JOIN ... ONconditions.#4855@mattfaltynIncludes tables referenced only by subqueries in
MERGE ONpredicates.#4767@mattfaltynIncludes tables read by subqueries in
UPDATE SETassignment expressions.#4788@mattfaltynIncludes tables read by scalar subqueries inside
VALUESrows.#4852@mattfaltynRetains source datasets when a Snowflake
PIVOTwraps a derived table.v1.52.0Compare Source
Added
#4682@kacpermudaAdds one JSON context payload for parent-run info, replacing 6 separate config keys in dbt and Spark.
#4716@himakolavennuCaptures incremental model settings, such as strategy and partitioning, plus a run-wide full_refresh flag.
#4705@himakolavennuAdds exposures, such as dashboards and notebooks, as a dataset facet, so lineage now includes table-to-exposure links.
#4744@wangxiaojingAdds a setting to turn off checkpoint tracking, cutting unnecessary RUNNING events and REST calls.
#4698@maitraymukeshkumarmodi-aimlImplements file size checks, so file size expectation results now produce a fileSize facet instead of being ignored.
#4680@kacpermudaAdds an optional facets field to ParentRunFacet, so producers can forward parent and root facets to child events.
#4711@HeroCCAdds the new UK1 and US2 FedRAMP Datadog sites to the transport's accepted site list.
Fixed
#4743@zerafachrisFixes a crash when a facet field is named additionalProperties, which confused the JSON deserialiser.
#4731@chuenchen309Fixes a crash when merging a scalar config value with a dict value from another source.
#4728@mattfaltynFixes shutdown so it waits for every event, even when several events share the same ID.
#4694@mobuchowskiFixes high CPU use and a memory leak in the Datadog transport's async worker.
#4729@zerafachrisFixes a pickling error for facets built with with_additional_properties().
#4722@chuenchen309Fixes tag matching so a user tag overrides an integration tag, regardless of letter case.
#4696@hcthakur2004Restores the HTTP debug level after a failed emit, so it does not stay switched on for later requests.
#4730@chuenchen309Fixes job name handling so
--openlineage-dbt-job-name=valueno longer breaks the dbt run.#4725@chuenchen309Fixes the profiles directory lookup so it now falls back to
~/.dbt/as intended.#4732@chuenchen309Fixes seed input resolution so tests use the seed's alias instead of its logical name.
#4724@chuenchen309Fixes a missing colon in the default Spark port, which broke the dataset namespace.
#4723@chuenchen309Fixes the SQL query externalQuery facet to use the dataset namespace, not a placeholder value.
#4752@mattfaltynIncludes tables referenced only in a WHERE subquery, such as WHERE EXISTS, in table lineage.
v1.51.0Compare Source
Added
#4653@mobuchowskiCapture per-model
configvalues (e.g.materialized,access,owner,group) and the user-definedmetamap from the dbt manifest into a newdbt_modeldataset facet, attached in both the legacy/local and structured-logs dbt processors.#3674@arturowczarekAdd a catalog handler for the open source Unity Catalog's
UCSingleCatalogSpark catalog so output dataset facets are produced when using OSS Unity Catalog, which cannot reuse the existingDeltaHandlerimplementation.#4592@jakub-moravecExtend the Job Type facet so producers, particularly streaming producers that emit a single event covering a time period rather than START/RUNNING/COMPLETE events, can document the lifecycle of the OpenLineage events consumers should expect.
Fixed
#4676@mishrasangeeta87Fix a bug where
df.write.mode("append").insertInto(table)did not emit an output dataset, by extracting the target table directly from theAppendDatalogical plan, aligning append-mode handling with overwrite-mode.#4660@mishrasangeeta87Build Databricks Unity Catalog table symlinks using the fully qualified table name and a
unity-catalognamespace instead of one derived from the underlying storage path, giving lineage consumers stable, catalog-qualified identifiers.#4652@tnazarewRelocate shaded classes per-transport (e.g.
io.openlineage.client.transports.gcs.shaded.x) instead of to a single shared shaded package, fixing Google library version conflicts when combininggcsandgcplineagetransports in a composite transport.#4661@arturowczarekTreat
abfs://andwasb://URIs the same as their TLS variants (abfss/wasbs) when deriving the object-storage namespace, since they point at the same underlying files and only differ in transport.#4666@mobuchowskiPreserve the underlying exception instead of only its message when a
CompositeTransportemit fails, so the real cause is visible instead of being lost.#4624@hcthakur2004Avoid indexing past the end of command-line arguments when a trailing option has no value.
#4627@hcthakur2004Avoid mutating the caller-provided chain in
get_from_nullable_chain, keeping it reusable after lookup.#4679@hcthakur2004Preserve the port on a datasource URL when Great Expectations strips credentials from it.
#4586@hcthakur2004Replace shared, mutable Datadog config dictionary defaults with per-instance factories so retry and async transport rule defaults aren't accidentally shared across instances.
#4656@hcthakur2004Normalize local
file://URIs before opening file transport paths so events can be appended through afile://log path.#4665@arturowczarekDetect Iceberg by the presence of the
iceberg-sparkSparkCatalogclass rather thaniceberg-core'sCatalog, so the Iceberg handler no longer activates in environments whereiceberg-coreis on the classpath withouticeberg-spark.v1.50.0Compare Source
Added
#4613@jsingh-yelpAdd a timeout-aware
emitoverload to the Java client andTransportAPI, allowing callers to provide a bounded wait when emittingRunEvents; Kafka transport uses the timeout to wait for producer acknowledgement before returning.#4596@jsingh-yelpEnable full OpenLineage lifecycle events (RUNNING, COMPLETE, FAIL) for Flink jobs submitted in detached Session Mode by introducing
OpenLineageDetachedJobStatusChangedListenerthat initialises the Flink job ID from the REST API, bypassing the missingJobCreatedEventon the JobManager side.Fixed
#4615@wangxiaojingMap
JobStatus.CANCELEDtoEventType.ABORTin Flink 2 job status handling; previously all non-FINISHED terminal statuses were mapped to FAIL, causing user-canceled jobs to appear as failures in downstream OpenLineage consumers.#4635@kacpermudaMerge client-configured environment variables into any
environmentVariablesfacet already set by the producer instead of replacing it; event-supplied values take precedence and a warning is logged on conflict.#4633@fm100Fix missing
IcebergScanReportinput dataset facet in CTAS/RTAS queries by reading the report directly fromscanReportSupplierwhen available, rather than relying onOpenLineageMetricsReporterwhich is registered after the scan has already occurred.v1.49.0Compare Source
Added
#4610@matveeysvAdd
CassandraJdbcExtractorto parse Cassandra JDBC URLs according to the driver specification, enabling lineage tracking for Cassandra databases via JDBC.Fixed
#4599@fm100Fix missing column lineage facet in
DbtStructuredLogsProcessorwhen using--consume-structured-logsoption by attaching the column lineage facet to the output dataset on node finished events.#4591@fm100Use a fully qualified job ID (with project ID and location) for the
externalQueryIdin theexternalQueryrun facet when using the dbt-bigquery adapter; also addsexternalQueryrun facet support when using--consume-structured-logs.#4598@codelixirExclude the
dataproc_job_attempt_timestamptag prefix when reading the job ID from Yarn tags in the GCP Dataproc facet, preventing the attempt timestamp from being reported as the job ID on retried jobs.#4611@tnazarewFix incorrect Spark configuration properties used in the Lakehouse Hive Catalog detection logic and extend catalog detection to
V2SessionCatalogHandlerused by DatasetBuilders.#4602@mrpalash-amzFix missing column lineage when the Snowflake Spark connector uses quoted identifiers (e.g. in AWS Glue ETL with
dbtablepath) by applyingstripQuotes()normalization on both sides of identifier comparisons inColumnLevelLineageBuilderandSqlCollector.v1.48.0Compare Source
Added
#4558@mobuchowskiFix incorrect Glue table symlink attachment when using S3 Tables as the Iceberg catalog by adding dedicated S3 Tables catalog support and preventing Glue fallback when the S3 Tables catalog is active.
#4574@tnazarewAdd support for the Lakehouse (formerly BigLake) Hive Catalog used by Spark jobs running on Managed Spark (formerly Dataproc).
#4546@adnanhemaniAdd support for the Snowflake Horizon Iceberg REST catalog, enabling the Spark integration to emit correct dataset identifiers with Snowflake namespaces for Iceberg tables managed through Snowflake's REST catalog.
Changed
#4536@W-ElyUpgrade the Amazon Kinesis Producer Library (KPL) from 0.15.12 to 1.0.7 and update AWS SDK dependencies to 2.44.3 to maintain compatibility and pick up upstream fixes.
#4552@kacpermudaAdd a warning log message when JSON parsing of OpenLineage configuration from environment variables fails, making configuration errors easier to diagnose.
#4557@mobuchowskiChange
BaseCatalogTypeHandler.getIdentifier()to returnOptional<DatasetIdentifier>instead of a nullable value, making the contract for Iceberg catalog type handlers more explicit and null-safe.#4587@tnazarewRefactor Spark integration's dataset-building logic from overloaded methods to a more flexible builder pattern, improving extensibility and maintainability.
Fixed
#4561@hcthakur2004Fix pull request number detection in the Python client to recognize
refs/pull/<number>/headformat (used in GitHub Actionspull_requestevents) in addition to the existingrefs/pull/<number>/mergeform.#4571@mobuchowskiFix
SnowflakeCatalogTypeHandler.getIdentifier()to returnOptional<DatasetIdentifier>as required byBaseCatalogTypeHandler, correcting a type mismatch introduced when the Snowflake Iceberg REST catalog handler was first added.#4556@dbathie-wtgMap Spark JDBC
sqlserverURLs toMsSqlDialectso SQL Server queries using bracketed identifiers ([schema].[table]) are parsed correctly during lineage extraction.v1.47.1Compare Source
Changed
#4501@tnazarewReplace the
quicktype-based text-manipulation generator with a structured pipeline that parses spec files viago-jsonschema, resolves references, and renders facet classes — enabling extension of generated classes (e.g. byool resources). Also renamesRun→RunWrapperandRunInfo→Runfor consistency withJobandDataset.Fixed
#4542@mobuchowskiFix
extract_adapter_typelookup for the Microsoft Fabric adapter by aligning theAdapterenum name with the dbt adapter type (fabric), while keeping thefabric-warehouseOpenLineage namespace unchanged.v1.47.0Compare Source
Added
#4485@harelsRegister
dbt-fabricin the supportedAdapterenum and addfabric://namespace extraction, preventingNotImplementedErroron Fabric profiles and enabling lineage tracking for Microsoft Fabric datasets.#4495@och5351Add
ClickHouseJdbcExtractorto enable lineage tracking for ClickHouse JDBC connections, supporting bothjdbc:clickhouse://andjdbc:ch://URL schemes with optional protocol prefix removal.#4506@jsingh-yelpAdd
ConfigFacetVisitorandSchemaFacetVisitorfor non-Table API Flink 2 datasets, enabling schema and configuration metadata extraction for DataStream API connectors beyond the Table API.#4457@kacpermudaExtend
DataQualityAssertionsDatasetFacetwith additional fields (actual value, expected value, severity) aligned withTestRunFacet, bumping the facet spec to 1-1-0.Changed
#4500@mobuchowskiPopulate
actualandexpectedfields onDataQualityAssertionsDatasetFacetandTestRunFacetwith the dbt test failure count and error threshold, enabling consumers to distinguish passing tests from failures and understand configured tolerances.Fixed
#4497@och5351Fix MySQL JDBC extractor to not prepend the URL database to already-qualified table names, preventing invalid 3-level identifiers like
mydb.schema.table1(MySQL treats DATABASE and SCHEMA as synonyms).#4520@mobuchowskiFix dbt-ol to recognize
retryas a valid dbt command so lineage is captured when re-running failed nodes.#4523@mobuchowskiFix aggregate test events in the legacy
consume_local_artifactspath that emitted COMPLETE (success) even when the embeddedDataQualityAssertionsfacet showed error-severity assertion failures.#4472@mobuchowskiFix stale
manifest.jsonfrom a prior invocation being loaded before dbt finishes its parse phase by subscribing to theArtifactWrittenstructured log event (dbt ≥ 1.9) and lazy-loading for older versions.#4489@hcthakur2004Fix
AttributeError: 'NoneType' object has no attribute 'startswith'crash whennode_info.unique_idis absent from the event by adding an earlyNoneguard.#4515@mobuchowskiFix crash when the dbt
ownermetadata field is a list rather than a string, handling both single-owner and multi-owner configurations gracefully.#4499@ichirotakamiFix
KeyErrorinextract_dataset_data()for dbt-fusion manifests where source nodes omit thedescriptionkey by using.get()with an empty string default.#4522@mobuchowskiFix
root=Nonein the parent facet for per-node events emitted via the legacyconsume_local_artifactspath by propagatingroot_parent_*fields todbt_run_metadata.#4503@hcthakur2004Fix file reading in the dbt provider to explicitly specify UTF-8 encoding, preventing
UnicodeDecodeErroron systems where the default locale encoding is not UTF-8.#4498@1fanwangFix
NullPointerExceptionwhen processing Hive queries with 3+ wayUNION ALL,INTERSECT, orEXCEPTset operations by recursively descending through intermediateQBExprnodes instead of assuming leaf-only structure.#4439@NETIZEN-11Fix
ClassCastExceptionwhen Iceberg returnsSparkChangelogTableinstead ofSparkTableby replacing the unsafe direct cast withinstanceofchecks and graceful fallback handling.#4521@Poojitha-R-RaoUpgrade
httpclient5dependency to 5.6.1 to address security vulnerability CVE-2026-40542.#4505@creazyfrogFix incorrect use of Hadoop's 3-argument
Pathconstructor that produced malformed S3 URIs likes3://bucket/prefix://database.db./table_name, which causedIllegalArgumentExceptionand silently dropped all lineage events.#4427@Poojitha-R-RaoUpgrade AWS SDK version to address security vulnerability CVE-2026-33871 (netty-codec).
v1.46.0Compare Source
Added
#4358@tnazarewAdd a new OpenLineage Go client with code generation from the spec, HTTP and GCP Lineage transport implementations, and CI integration.
#4398@himakolavennuAdd
sourceCodeLocationjob facet emission for dbt events, with repo URL configurable via--openlineage-repo-urlorOPENLINEAGE_REPO_URL, falling back togit remote get-url originautodetection. URLs are normalized tohost/org/repoformat.#4449@mobuchowskiEmit
TestRunFacetfrom both the regular dbt processor path and the structured logs path for test result nodes, including outcome (pass/fail/warn/error), severity, and expected/actual values.#4436@himakolavennuAdd optional
pullRequestNumberfield toSourceCodeLocationJobFacet(spec bumped to 1-1-0) with auto-detection from CI environment variables (GITHUB_REFfor GitHub Actions,CI_MERGE_REQUEST_IIDfor GitLab CI).#4448@mobuchowskiAdd
TestRunFacetJSON Schema for recording test outcomes (pass/fail/warn/error), severity, expected vs actual values, and test message, along with generated Python client code and facet registration.Changed
#4403@himakolavennuMake git autodetect for
sourceCodeLocationopt-in by default (disabled=True) to avoid surprising subprocess calls during event emission, and reduce subprocess count from 4 to 2 by consolidating git calls.#4391@mobuchowskiRemove support for Python 3.9, which reached end-of-life in November 2024. The minimum supported Python version is now 3.10.
Fixed
#4451@bahram-cdtFix broken SPI directory path (
META-INF.services/→META-INF/services/), wrong class name in SPI file (missing.kinesissub-package), and NullPointerException whenpropertiesconfig is omitted in the Kinesis transport.#4400@mobuchowskiFix the structured logs processor to correctly resolve parent datasets for dbt singular tests using
parent_map, so test result assertions are now attached to the correct datasets asDataQualityAssertionsDatasetFacet.#4390@orthoxeroxFix column-level lineage for InMemoryRelation when it contains non-unique unqualified column names by using
qualifiedNameinstead ofnamefor correct column mapping.#4384@usamakunwarFix dataset symlinks when using Glue catalog for Iceberg by detecting Glue-based catalogs via Spark configuration and generating correct ARNs even when not using the Hive metastore.
#4366@mobuchowskiFix JDBC column lineage to skip
extractInternalInputswhen SQL-based column lineage is already available, preventing alias names from being incorrectly included as input fields alongside the original column names.v1.45.0Compare Source
Added
#4368@jsingh-yelpAdd support for emitting DatasetConfigFacet in the Flink native listener, enabling configuration tracking for datasets processed by Flink jobs.
#3747@dolfinusIntroduce the HierarchyDatasetFacet to provide structured representation of dataset hierarchy levels (database, schema, table, etc.) without relying on dataset name parsing, enabling consistent handling across different database systems with varying hierarchy depths.
Changed
#4383@mobuchowskiImprove performance by switching from expensive semanticHash() calls to identity-based tracking using IdentityHashMap, eliminating hot path in large jobs during plan traversal.
#4376@mobuchowskiOptimize findDependentInputs by replacing LinkedList with HashSet for visited node tracking, improving lookup performance from O(n) to O(1).
Fixed
#4372@mobuchowskiFix handling of dbt singular tests when processing structured logs to ensure proper test result tracking and lineage extraction.
#4357@tnazarewFix the condition for adding project_id to catalog properties, ensuring proper BigLake catalog detection and handling.
v1.44.1Compare Source
Fixed
#4349@harelsAttach
ExtractionErrorRunFacetto run events when@handle_keyerror-decorated extraction methods fail, making previously invisible extraction errors visible to downstream consumers instead of silently emitting incomplete events._get_model_node#4348@harelsFix exception type mismatch in
_get_model_node()by raisingKeyErrorinstead ofRuntimeError, allowing the@handle_keyerrordecorator to catch it and returnNonegracefully when a node_id is not found in the manifest..get()for optional project version retrieval#4345@zagoodmanFix crash when
versionkey is absent fromdbt_project.yml, which became optional in dbt 1.5, by using.get()instead of direct key access.v1.44.0Compare Source
Added
#4313@jakub-moravecAdd JWT authenticator for Java and Python clients, enabling token-based authentication without requiring a custom authenticator implementation.
#4283@kchledowskiEnable extraction of input dataset symlinks from DataSourceRDD, providing richer lineage information for RDD-based Iceberg operations.
Changed
#4329@kchledowskiDisable column-level lineage extraction for LogicalRDD plans to prevent incorrect lineage caused by lost schema and transformation context.
#4331@kchledowskiDisable unreliable input schema extraction from LogicalRDD and instead extract schemas from Iceberg table metadata when reading via DataSourceRDD.
Fixed
#4285@LegendPawel-MarutAlign schema definitions for dbt-run-run-facet and dbt-version-run-facet to fix validation inconsistencies.
#4320@ah12068*Handle missing profiles_dir key in
Configuration
📅 Schedule: (UTC)
🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.
♻ Rebasing: Whenever PR is behind base branch, or you tick the rebase/retry checkbox.
🔕 Ignore: Close this PR and you won't be reminded about this update again.
This PR was generated by Mend Renovate. View the repository job log.