Skip to content

Integrating ClickHouse with Google Meet

The Google Meet registry item copies raw readers with fixed discovery windows and journaled conference, transcript, and entry progress into a chkit project.

Terminal window
bunx chkit add google-meet --with-tests
bunx chkit check
bunx chkit generate --name add_google_meet
bunx chkit migrate --apply
bunx chkit ingest run --tag provider:google-meet
bunx chkit ingest status --tag provider:google-meet

Set GOOGLE_MEET_ACCESS_TOKEN with meetings.space.readonly access before ingestion. Review source identity, window settings, schema placement, and limits in src/integrations/google-meet/config.ts. Keep the source identity tied to one Google account; these list responses do not identify the authenticated account. Schema imports do not call Google.

ResourceDefault ClickHouse tableRecords syncedAPI reference
Conferences (conferences)google_meet_conferences_rawConference records visible to the access tokenGET/conferenceRecords
Transcripts (transcripts)google_meet_transcripts_rawTranscript metadata for discovered and retained pending conferencesGET/conferenceRecords/{conferenceRecord}/transcripts
Transcript entries (transcript-entries)google_meet_transcript_entries_rawIndividually keyed, paged transcript entries for retained conferencesGET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/transcripts, GET/conferenceRecords/{conferenceRecord}/transcripts/{transcript}/entries

One installation exports one pipeline with independently selected conferences, transcripts, and transcript-entries streams. Each owns its journal progress. Transcript and entry streams discover their own parent conferences, so they work without running the conferences stream first. Nested pagination follows parent-scoped API endpoints inside the selected resource stream.

Terminal window
bunx chkit ingest run --tag provider:google-meet --tag resource:conferences
bunx chkit ingest run --tag provider:google-meet --tag resource:transcripts
bunx chkit ingest run --tag provider:google-meet --tag resource:transcript-entries

pipeline.ts wires destinations and resource streams; sources/resources.ts shares the provider state machine, and client.ts handles requests. sourceId scopes raw rows and query state; streamPrefix determines journal IDs. Defaults preserve google-meet.* checkpoints. Another installation needs distinct values for both. The pipeline factory binds configuration and injectable request dependencies while using the exported destinations. Edit database in config.ts before schema imports to change table placement.

The default pipeline runs one stream and request at a time. Filtered runs give resources separate schedules and execution budgets.

Each stream independently discovers completed conferences through fixed end_time windows and ongoing conferences through end_time IS NULL. End-time discovery includes long-running calls that began before the lookback. The initial range defaults to 30 days, later discovery overlaps seven days, and each cycle advances by at most windowDays.

Transcript and entry streams retain parent conferences until provider expiry and revisit them on completed cycles. An empty transcript list or a still-processing artifact stays eligible after discovery advances, so late transcripts are collected outside the discovery overlap. This is polling; Meet provides no modification/change token and no transcript or entry date filter. See conference filters and artifact behavior.

The journal stores fixed cycle bounds and nested parent, transcript-page, and entry-page positions. Candidate progress commits after its covering rows load; failed loads retain the preceding position. Rerun after interruption or budget exhaustion to continue. Rejected page tokens replay their scoped collection once, and repeated or malformed tokens fail visibly.

Saved recovery counts cover each unfinished discovery phase, parent transcript collection, and transcript entry collection. They clear at the acknowledged terminal boundary, so pauses and nonterminal replay pages do not renew the allowance. A second rejection fails across executions. Adjust maxChunks in config.ts, execution duration, or polling frequency, then review coverage before explicitly migrating saved state or using a new stream identity for a fresh scan.

Streams permit 200 chunks per run and retain at most 1,000 conferences. Edit maxChunks and maxPendingConferences in config.ts; an oversized pending queue fails instead of losing parents. Source/window changes require explicit checkpoint migration or a new stream identity. Native JSON requires ClickHouse 25.3 or later. Schedule repeat runs externally with one ingestion process per destination at a time.

Google removes conference records and API transcript entries 30 days after a call ends. Expiry during unfinished work fails with a coverage-gap message. Previously checked parents leave the queue on expiry, while stored raw rows remain. Expired or inaccessible resources do not become inferred deletions. See conference retention and artifact retention.

Raw identities are [sourceId, provider.name]. Transcript metadata includes conference_name; individual entry rows include conference_name and transcript_name. Entries stream page by page into google_meet_transcript_entries_raw. The API entries do not capture later edits to the separate Google Docs transcript file.

Version 0.2.0 adds the entry destination and stores new transcript metadata without embedding every entry in one row. Existing tables remain available. Row identities now include the source label; keep older datasets as archives or migrate their identities deliberately. See the installed README for upgrade details.

Explicit backfill bounds constrain conference end times and skip ongoing-call discovery. Reuse the same backfill ID and bounds to resume; expired API history remains unavailable. Artifact lists are parent-scoped and cannot themselves be date-filtered.

Terminal window
bunx chkit ingest run --tag provider:google-meet --backfill october --from 2026-10-01T00:00:00Z --to 2026-10-05T00:00:00Z
Terminal window
bun test src/integrations/google-meet/tests/basic.test.ts

Fixtures cover nested entry resumption, late transcripts from retained parents, fixed end-time and ongoing discovery, rejected tokens, sink failures, isolated backfill bounds, and expiry gaps.

Version 0.2.0

  • Checkpoint fixed conference discovery windows and retain pending conferences until expiry to revisit late transcript artifacts.
  • Resume acknowledged discovery, transcript, and entry pages with bounded token recovery and visible expiry gaps.
  • Store transcript entries in a separate raw table and scope row IDs to the source; migrate legacy identities deliberately when upgrading.
  • Separate editable installation configuration, injectable HTTP clients, and resource readers from one pipeline; preserve independent stream identities and provider recovery.

Version 0.1.0

  • Introduce raw conference and transcript ingestion with transcript entries nested in each transcript observation.