Integrating ClickHouse with GitHub
The GitHub registry item copies a small API reader and raw ClickHouse tables into a chkit project.
Install
Section titled “Install”bunx chkit add githubbunx chkit checkbunx chkit generate --name add_githubbunx chkit migrate --applybunx chkit ingest run --tag provider:githubSet GITHUB_TOKEN in the runtime environment before the ingestion run. Edit githubConfig.repositories in src/integrations/github/config.ts to select repositories. pipeline.ts groups independent issues and stargazers streams for each repository in one pipeline; sources/ contains the resource readers and raw table declarations. The named pipeline and table exports remain in index.ts.
The pipeline factory binds source configuration and injectable HTTP dependencies. Keep sourceId stable for one installation; its stream IDs own the checkpoints. Table schemas are edited in their source modules before migration.
| Resource | Default ClickHouse table | Records synced | API reference |
|---|---|---|---|
Issues and pull requests (issues) | github_issues_raw | Issues and pull requests updated inside a timestamp window, including paginated comments and pull request context. | GET/repos/{owner}/{repo}/issues |
Stargazers (stargazers) | github_stargazers_raw | Repository stargazers with starred_at when the token permits it | GET/repos/{owner}/{repo}/stargazers |
Sync behavior
Section titled “Sync behavior”Issues use since with update-time ordering, a fixed local upper cutoff, and five minutes of overlap. GitHub has no server-side upper timestamp filter. The checkpoint advances only after the complete issue window and destination writes succeed; failed windows replay from their lower bound. The reader follows provider Link URLs and includes paginated issue comments plus pull request detail, reviews, review comments, commits and files. Provider caps are reflected in raw _chkit_context completeness fields.
Stargazers use the built-in fullSync strategy; the journal records completion even for an empty read. Both readers use paginate for request retries and continuation cycle detection. Removed stars and deleted issues remain stored. Live pagination and child-only updates require periodic reconciliation or a durable webhook companion. GitHub issue filters and pagination describe the API behavior.
bunx chkit ingest run --tag provider:github --tag resource:issues --backfill reconcile-2026-10-04 --from 2008-01-01bunx chkit ingest status --tag provider:github --jsonHistorical replay reads current observations, not historical versions. Use a new backfill ID for each full reconciliation. The raw tables preserve provider fields and require ClickHouse 25.3 or later. See the installed README for token permissions and registry installation.
Select a repository/resource stream independently:
bunx chkit ingest run --tag provider:github --tag repository:obsessiondb/chkit --tag resource:issuesThe default pipeline executes streams sequentially with maxStreams: 1. Separate filtered runs give each stream its own schedule and execution budget; a failed stream does not advance another stream’s progress.
Test the reader
Section titled “Test the reader”bunx chkit add github --with-testsbun test src/integrations/github/tests/basic.test.tsChangelog
Section titled “Changelog”Version 0.2.0
- Sync issues through overlapping updated-time windows, advancing the watermark only after complete destination acknowledgement.
- Collect all accessible issue comments and pull-request context, reporting provider caps on commits and files.
- Group independent repository/resource streams in one configurable pipeline, using paginate and fullSync for stargazers without custom scan state.
- Separate editable configuration, injectable HTTP client, and resource readers; cover isolated selection, retries, and failure recovery without changing raw destinations.
Version 0.1.0
- Introduce raw repository issues and stargazers ingestion.