Pipeline architecture
Data flows through three stages: Databricks Unity Catalog → S3 → RDS PostgreSQL. The API has no direct connection to Databricks at runtime — it only ever reads from RDS. The underlying Databricks jobs are namedparquet_to_s3 and ingest_parquet_to_rds.
Trassets’ internal documentation describes this sync as running daily. This
repository’s public API layer has no machine-enforced schedule of its own to verify
against — the sync cadence is a property of the upstream pipeline, not something the API
configures or exposes. Do not assume sub-daily freshness; always check
last_sync_at (see
below) rather than assuming a fixed interval has elapsed.TRUNCATE + Reload semantics
On every sync run, each table is dropped and reloaded in full — there is no row-level incremental update at the table level. This has two direct consequences for integrators:- No partial sync. A table is either fully replaced by the latest sync, or it isn’t
touched at all. There is no mechanism to sync “only the rows that changed” within a
single table —
GET /changes(below) tells you which tables changed, not which rows. - Schema changes are always breaking for cumulative tables. For tables that
accumulate history across reloads (for example
contract_ends), a schema change cannot be applied additively — TRUNCATE + Reload always replaces the entire table content, so there is no way to migrate old rows to a new shape in place. Any schema change to such a table is treated as a breaking change requiring a new API version, with both the old and new table shapes synced in parallel during a deprecation window.
When last_sync_at is set
Sync state is tracked per (customer_id, table_name) in an internal status table with
these fields: last_sync_at (nullable timestamp), status (synced, pending, or
error), row_count, primary_key_column, and error_message.
last_sync_at is set only when a sync completes successfully for that table — it stays
NULL until the first successful sync, and a failed sync (status = 'error') does not
advance it. This means:
- A table with
status = 'pending'orstatus = 'error'will never appear inGET /changes, regardless of how recently a sync was attempted. last_sync_atreflects table-level completion, not the modification time of any individual row inside that table.
What GET /changes actually checks
GET /changes?since=<timestamp> returns a table only if all of the following hold:
statusis exactly'synced'(not any other status value)last_sync_atis notNULLlast_sync_at > since— a strict greater-than comparison, not>=
ETag response header is the maximum last_sync_at across all of a customer’s synced
tables, so polling with If-None-Match lets you skip a request entirely when nothing has
synced since your last poll (see Delta Sync for the full polling
pattern).
Practical implications for your integration
- Poll
/changes, don’t guess a schedule. Because there is no documented, enforced sync interval, build your integration to react to what/changesreports rather than assuming data updates at a fixed time of day. - A changed table means “re-fetch everything for this table,” not “here are the diffs.”
Row-level change detection does not exist yet — when a table appears in
/changes, the safest action is a full re-fetch of that category or raw endpoint. - Treat
last_sync_atas a freshness signal, not a change log. It tells you the table was successfully reloaded at that time; it does not tell you which rows differ from the previous version.