FastSync
FastSync bypasses Singer JSON for supported bulk transfers and uses native database export, staging, copy, and merge operations. It has two components:
Component |
Selection |
Target effect |
|---|---|---|
FullSync |
Initial load, |
Publishes a complete source-table copy. |
PartialSync |
|
Merges a filtered source range into the existing target. |
FastSync is not a replication method. LOG_BASED, INCREMENTAL, and
FULL_TABLE remain the table’s replication methods.
Supported routes
Source |
Target |
Table formats |
FullSync |
PartialSync |
|---|---|---|---|---|
MariaDB / MySQL |
Snowflake |
Native, managed Iceberg v3 |
Yes |
Yes |
PostgreSQL |
Snowflake |
Native, managed Iceberg v3 |
Yes |
Yes |
MongoDB |
Snowflake |
Native |
Yes |
No |
MariaDB / MySQL |
PostgreSQL |
Native |
Yes |
No |
PostgreSQL |
PostgreSQL |
Native |
Yes |
No |
MongoDB |
PostgreSQL |
Native |
Yes |
No |
Endpoint support status from Connectors still applies. For a
route without a FastSync component, a normal run_tap falls back to Singer.
The explicit fast_sync command instead fails without loading data.
Automatic selection and handover
During run_tap, PipelineWise selects FullSync when:
the table uses
FULL_TABLE;an
INCREMENTALtable has no replication-key bookmark; ora
LOG_BASEDtable has no LSN, binlog, GTID, or change-stream bookmark.
After a successful initial FullSync, PipelineWise writes the captured bookmark
and starts the Singer portion of the same run_tap invocation for incremental
or log-based tables. It does not wait for the next scheduled launch.
If FullSync fails, Singer does not advance that table past an incomplete initial load. Restart the same command after correcting the failure.
Snowflake Iceberg publication
MariaDB/MySQL and PostgreSQL FastSync use native staging for managed Iceberg v3.
A missing table is created through explicit-schema CTAS. An exactly compatible
FullSync target uses INSERT OVERWRITE; a compatible new nullable column is
added first. Other FullSync mismatches require guarded replacement. PartialSync
uses transactional range DML and requires a primary key.
The route requires explicit target_table_format: iceberg,
iceberg_version: 3, and data_flattening_max_level: 0. Native remains the
default. See Snowflake Iceberg tables for metadata limits, writer exclusion, and
recovery.
After an eligible initial load, Singer continues LOG_BASED or INCREMENTAL
replication against the same managed-v3 table in the same run.
FastSync availability and its zero-flattening requirement are specific to these
routes. A compatible Singer-only source such as Salesforce can load managed v3
through target-snowflake without gaining a FastSync component; it retains
its normal flattening setting and sends FULL_TABLE streams through Singer.
Snowflake source-side transformations
PostgreSQL and MySQL/MariaDB FullSync and PartialSync apply configured top-level
transformations in the source SELECT, before generating CSV. Native and
managed Iceberg v3 staging, S3 uploads, and archived load files contain only the
configured transformed output. Unsupported rules fail before export rather than
falling back to Snowflake updates. Source range selection and replication
bookmarks remain unchanged. See Load-time transformations for execution order,
conditional rules, supported SQL semantics, and upgrade recovery requirements.
Snowflake string widths
MariaDB/MySQL and PostgreSQL FastSync declare string staging and new-target
columns as VARCHAR(134217728) for both native and managed Iceberg v3 routes.
This applies to character and text families, MariaDB/MySQL blob and enum types,
and the fallback for an otherwise unmapped source type. It avoids Snowflake’s
narrower 16,777,216-character default for a bare VARCHAR.
The declared limit does not remove Snowflake’s 128 MB encoded-value limit, so a value can reach the byte limit before the character limit when it contains multi-byte characters. Values above Snowflake’s maximum still fail while loading the native staging table, before publication or state advancement. See Snowflake’s string and binary data type reference.
An existing native PartialSync target whose compatible text column is narrower is widened before the merge. Existing managed Iceberg v3 string columns must already use the maximum width; see Snowflake Iceberg tables.
Singer schema evolution has a different existing-native policy:
target-snowflake uses the maximum width for new string columns but does not
widen or version a compatible existing native string column solely because of
its declared width. See Snowflake target.
MariaDB/MySQL and PostgreSQL Snowflake FullSync and PartialSync preserve LF, CR,
CRLF, tab, CSV punctuation, and literal backslash sequences in string values.
PostgreSQL preserves Unicode; MariaDB/MySQL uses an utf8mb4 projection and
default connection encoding, including supplementary Unicode, while continuing
to remove NUL characters. This preservation applies to native and
managed Iceberg v3 publication; the table format does not change the staged
string representation. SQL NULL and a literal \N string remain distinct.
This behavior does not repair values normalized by an earlier run. Resync each affected table from its source after upgrading.
Explicit FullSync
fast_sync resyncs every selected table regardless of its current bookmark:
pipelinewise fast_sync \
--tap <tap_id> \
--target <target_id> \
--tables <schema.table>
This operation can replace target data and reset replication bookmarks. Review Resync and repair before running it against a large or actively written table.
An unfiltered PostgreSQL fast_sync containing LOG_BASED tables resets the
tap-specific slot once before workers start, with or without --force.
Filtered FastSync, automatic initial loads, and standalone PartialSync retain
it. --force only bypasses the resync size limit; configured
sync_start_from ranges remain in effect. See PostgreSQL source-slot reset
for exact commands, safety checks, pending-Iceberg guards, and failure recovery.
Configured PartialSync
sync_start_from makes explicit fast_sync use PartialSync for that table:
tables:
- table_name: "orders"
replication_method: "LOG_BASED"
sync_start_from:
column: "updated_at"
static_value: "2024-01-01"
drop_target_table: false
Setting |
Required |
Behaviour |
|---|---|---|
|
Yes |
Applies |
|
Exactly one value source |
Uses the same literal boundary on every run. |
|
Exactly one value source |
Runs a source query that must return one row and one column. |
|
No; default |
Recreates the target before loading the filtered result. |
Exactly one of static_value and dynamic_value is allowed. Configured
PartialSync is supported only from MariaDB/MySQL or PostgreSQL to Snowflake.
If a dynamic query returns no boundary, PipelineWise treats the partial range as empty and completes successfully. Use a static boundary when an empty result would hide a configuration or source-data problem.