Integrating ClickHouse with Lemlist
The Lemlist registry item copies a small API reader and raw ClickHouse tables into a chkit project.
Install
Section titled “Install”bunx chkit add lemlistbunx chkit checkbunx chkit generate --name add_lemlistbunx chkit migrate --applybunx chkit ingest run --tag provider:lemlistSet LEMLIST_API_KEY in the runtime environment before the ingestion run. Edit src/integrations/lemlist/config.ts for the account namespace, page size, initial date, overlap, and interval length. pipeline.ts groups independent activities and campaigns streams in one account pipeline; select either resource with tags. sources/ contains the readers and client.ts uses paginate() for authenticated requests. index.ts exports the pipeline and raw schema.
createLemlistPipeline(config, deps) binds reader settings and optional HTTP dependencies. Configure storage in the installation’s lemlistConfig before schema discovery; factories use the exported raw tables.
Keep each stream ID and destination tied to one account. The reader does not resolve the key’s account identity; switching accounts requires new stream IDs and destinations so previous watermarks cannot skip history.
| Resource | Default ClickHouse table | Records synced | API reference |
|---|---|---|---|
Activities (activities) | lemlist_activities_raw | Outreach activity events | GET/activities |
Campaigns (campaigns) | lemlist_campaigns_raw | Campaign metadata | GET/campaigns |
Sync behavior
Section titled “Sync behavior”Activities use inclusive minDate/maxDate filters on createdAt, thirty-day intervals, and one day of overlap. Each completed interval checkpoints only after its pages have destination acknowledgement. A failed interval starts at offset zero; an explicit backfill resumes its completed interval frontier in an isolated namespace. Empty intervals also commit progress. Offset pages remain mutable, so their offsets are never persisted as change cursors.
The activities API can backfill metadata on older activities. Periodic historical replay refreshes such changes and late arrivals beyond the overlap. Campaigns use creation-time ordering and ordinary fullSync() with journaled execution completion, including empty results. Failed scans replay from offset zero; creation-only incremental reads would miss campaign updates. Campaigns ignore date bounds and do not provide date-range backfills.
bunx chkit ingest run --tag provider:lemlist --tag resource:activities --backfill reconcile-2026-10-04 --from 2000-01-01bunx chkit ingest status --tag provider:lemlist --jsonUse a new backfill ID for each complete reconciliation; reusing an ID resumes that backfill. Narrowing a reused backfill’s upper bound below its committed frontier is rejected. Removed campaigns and activities remain stored. The raw tables preserve provider fields and require ClickHouse 25.3 or later. See the installed README and registry installation.
Test the reader
Section titled “Test the reader”bunx chkit add lemlist --with-testsbun test src/integrations/lemlist/tests/basic.test.tsChangelog
Section titled “Changelog”Version 0.2.0
- Sync activities through bounded, overlapping creation-time intervals and advance checkpoints only after all interval pages are acknowledged.
- Keep one account pipeline with independently selectable activities and campaigns streams, configurable readers, and shared pagination.
- Replay interrupted activity intervals from offset zero and use built-in full-sync completion for campaigns without treating creation time as a modification cursor.
- Preserve raw destinations and include fixtures for empty intervals, backfills, and source or destination failures.
Version 0.1.0
- Introduce raw campaigns and activities ingestion.