Compare commits

...
212 Commits
Author SHA1 Message Date
Max Deichmann 780880cf07 chore: release v2.89.0
CI/CD / lint (push) Waiting to run
CI/CD / test-docker-build (push) Waiting to run
CI/CD / tests-web (node20, pg12) (push) Waiting to run
CI/CD / tests-web (node20, pg15) (push) Waiting to run
CI/CD / tests-worker (node20, pg12) (push) Waiting to run
CI/CD / tests-worker (node20, pg15) (push) Waiting to run
CI/CD / e2e-tests (push) Waiting to run
CI/CD / e2e-server-tests (push) Waiting to run
CI/CD / all-ci-passed (push) Blocked by required conditions
CI/CD / push-docker-image (push) Blocked by required conditions
release.yml / release (push) Waiting to run
2024-11-12 19:34:07 +01:00
Max DeichmannandGitHub eb6f93bd9c fix: fix public trace access (#4195) 2024-11-12 18:32:45 +00:00
Max DeichmannandGitHub a0d4a2e599 fix: fix indefinite table queries (#4194) 2024-11-12 17:04:23 +00:00
Max DeichmannandGitHub 8458443a2e chore: remove logs (#4193) 2024-11-12 16:37:49 +01:00
Max DeichmannandGitHub feafc1b7d3 feat: move scores table apis to experimentation setup (#4192) 2024-11-12 15:14:33 +00:00
Max DeichmannandGitHub fcd54180bc chore: centralize lookback guarantees (#4188) 2024-11-12 15:56:33 +01:00
Max DeichmannandGitHub d73c3ec8d3 fix: fix clickhouse filters (#4190) 2024-11-12 14:48:45 +00:00
Max DeichmannandGitHub ca7210b26f feat: use timestamp to link single traces from navigation (#4187) 2024-11-12 11:50:51 +00:00
Max DeichmannandGitHub d9e22324d4 feat: up/down detail navigation support params (#4186) 2024-11-12 12:13:44 +01:00
Max DeichmannandGitHub 0a18e59274 perf: add timestamp to traces table link (#4185) 2024-11-12 10:38:00 +00:00
Max DeichmannandGitHub ffc5f81fb5 chore: add error logs in trpc route (#4184) 2024-11-12 09:56:03 +00:00
marliessophieandGitHub b4668ed851 Revert "fix(dataset_runs): ensure consistent ordering to fetch score data for dataset run aggregation metrics table" (#4182)
Revert "fix(dataset_runs): ensure consistent ordering to fetch score data for…"

This reverts commit 2b2022ae85.
2024-11-12 09:20:12 +00:00
Steffen SchmitzandGitHub 999ff06c58 chore: migrate spammy log messages to debug level (#4181) 2024-11-12 07:41:01 +00:00
Steffen SchmitzandGitHub e494b2c215 chore: remove uninstantiated projectId from clickhouse seeder (#4180) 2024-11-12 06:56:42 +00:00
marliessophieandGitHub 2b2022ae85 fix(dataset_runs): ensure consistent ordering to fetch score data for dataset run aggregation metrics table (#4176) 2024-11-11 16:07:56 +00:00
Steffen SchmitzandGitHub 3afa137130 chore: update latency calculation for clickhouse dashboard queries (#4175) 2024-11-11 16:28:43 +01:00
Max DeichmannandGitHub 0b88ddf438 fix: correctly filter observations by generations for generations ui table (#4173) 2024-11-11 14:16:01 +00:00
Marc KlingenandGitHub f878bd1317 perf(cloud): add 60 min stale time for usage indicator (#4171) 2024-11-11 14:10:58 +01:00
Marc KlingenandGitHub 72075b29de fix(ui): width of date range time input (#4170) 2024-11-11 14:06:58 +01:00
Max DeichmannandGitHub edf1a0b879 fix: fix aggregated user consumption (#4167) 2024-11-11 11:21:03 +00:00
Max DeichmannandGitHub df763ff09c perf: improve performance for latency dashboard (#4164) 2024-11-11 10:23:49 +00:00
Max DeichmannandGitHub 83e09bcf47 chore: read dashboards form clickhouse flag (#4162) 2024-11-10 20:24:50 +00:00
Max DeichmannandGitHub 10e9aef044 feat: move sessions api to experimentation setup (#4161) 2024-11-10 20:09:59 +00:00
Max DeichmannandGitHub 45a0c2eb3d feat: move generations api to experimentation setup (#4160)
chore: move generations api to experimentation setup
2024-11-10 19:09:44 +00:00
Max DeichmannandGitHub 99c2a75964 perf: add time filter for sessions (#4159) 2024-11-10 18:49:11 +00:00
Max DeichmannandGitHub 9cc61eb851 perf: add timestamp to filteroptions api (#4158) 2024-11-10 18:35:06 +00:00
Max DeichmannandGitHub 62e51b8cf6 feat: add traces experimentation (#4157) 2024-11-10 17:45:31 +00:00
Max DeichmannandGitHub b147560a80 feat: steer reading from CH and PG (#4153) 2024-11-10 18:14:30 +01:00
Max DeichmannandGitHub 97a0b1f44c perf: increase traces metrics performance (#4155) 2024-11-10 17:08:00 +01:00
Max DeichmannandGitHub 8c1e6647e0 feat: query model latencies from clickhouse (#4152) 2024-11-10 14:37:56 +00:00
Max DeichmannandGitHub 8432327127 fix: traces aggregate trace filter (#4151) 2024-11-10 14:12:29 +00:00
Max DeichmannandGitHub 231f8e8f0c feat: read latency tables from clickhouse (#4150) 2024-11-10 13:47:58 +00:00
Max DeichmannandGitHub 28e43eb74c feat: query scores over time (#4149) 2024-11-10 12:57:30 +00:00
Steffen SchmitzandGitHub e2d3647d76 chore: set source: API if no source provided (#4148) 2024-11-10 13:41:01 +01:00
Max DeichmannandGitHub 876509d9e5 feat: add user charts (#4146) 2024-11-10 12:18:38 +00:00
Max DeichmannandGitHub 1647a79198 perf: scores dashboard performance improvement (#4144)
fix
2024-11-10 12:01:16 +00:00
Steffen SchmitzandGitHub 622d8a89a7 chore: update bookmarked, tags, and public for clickhouse traces (#4132) 2024-11-10 11:29:48 +00:00
Max DeichmannandGitHub d5e66e10d6 feat: add table search for clickhouse (#4142) 2024-11-09 15:55:11 +01:00
Max DeichmannandGitHub 5b1776710f chore: clean ups of clickhouse read logic (#4141) 2024-11-09 14:35:04 +01:00
Max DeichmannandGitHub 7e7e3b77d4 fix: correctly parse model params from clickhouse (#4140) 2024-11-09 12:48:45 +00:00
Max DeichmannandGitHub 28ecd9600a fix: fix latency calculations when reading from clickhouse (#4139) 2024-11-09 10:59:52 +00:00
Max DeichmannandGitHub 1ef6834a31 feat: delete traces from clickhouse from the UI (#4138) 2024-11-09 11:49:44 +01:00
Marc KlingenandGitHub b88f3bed40 feat(ui): add notification card in sidebar (#4137) 2024-11-09 00:30:36 +00:00
Steffen SchmitzandGitHub eaa885f5f5 chore: inject eval and annotation scores into clickhouse (#4081) 2024-11-08 17:42:41 +00:00
Max DeichmannandGitHub f7f48c2509 fix: reduce ingestion concurrency (#4130) 2024-11-08 18:18:40 +01:00
Steffen SchmitzandGitHub 07344e89b1 chore: move legacy api endpoints to async ingestion (#4109) 2024-11-08 18:00:59 +01:00
Max DeichmannandGitHub e6bc717612 perf: improve generations table on clickhouse (#4124) 2024-11-08 11:18:41 +00:00
Max DeichmannandGitHub 11521fc24d perf: improve generations performance (#4111) 2024-11-07 20:54:49 +00:00
Max DeichmannandGitHub 213ec6027d perf: add sessions index in clickhouse (#4108) 2024-11-07 18:09:41 +01:00
Max DeichmannandGitHub 259d832d37 feat: support sessions API from Clickhouse (#4065) 2024-11-07 16:52:07 +00:00
Marc KlingenandGitHub 62deb83520 chore(cloud): stripe checkout settings (#4107) 2024-11-07 16:34:07 +00:00
Max DeichmannandGitHub edc4cd7873 perf: query traces via timestamps (#4105) 2024-11-07 14:52:33 +01:00
Steffen SchmitzandGitHub 05b75fbee0 chore: merge observation backfill states in table (#4104) 2024-11-07 13:20:18 +00:00
Steffen SchmitzandGitHub d6065e8b89 chore: add stateSuffix to observation migration for potential parallelization (#4101) 2024-11-07 12:46:58 +00:00
Steffen SchmitzandGitHub 617d129542 chore: swap negative provided usage for null (#4100) 2024-11-07 12:29:44 +00:00
Max Deichmann 057424b269 chore: release v2.88.0
CI/CD / lint (push) Waiting to run
CI/CD / test-docker-build (push) Waiting to run
CI/CD / tests-web (node20, pg12) (push) Waiting to run
CI/CD / tests-web (node20, pg15) (push) Waiting to run
CI/CD / tests-worker (node20, pg12) (push) Waiting to run
CI/CD / tests-worker (node20, pg15) (push) Waiting to run
CI/CD / e2e-tests (push) Waiting to run
CI/CD / e2e-server-tests (push) Waiting to run
CI/CD / all-ci-passed (push) Blocked by required conditions
CI/CD / push-docker-image (push) Blocked by required conditions
release.yml / release (push) Waiting to run
2024-11-07 11:44:46 +01:00
Max DeichmannandGitHub 63d1b85a6c fix: pagination for clickhouse tables (#4098) 2024-11-07 11:41:46 +01:00
Max DeichmannandGitHub 9481c3b424 feat: build scores table for clickhouse (#4092) 2024-11-07 10:18:14 +00:00
Steffen SchmitzandGitHub 15e174fc2f chore: convert starttime filter to DateTime in Ingestion pipeline (#4095) 2024-11-07 09:51:35 +01:00
Steffen SchmitzandGitHub 5efd4ec0a4 chore: include backgroundmigration state default in schema.prisma (#4094) 2024-11-07 08:25:32 +00:00
marliessophieandGitHub 895b61fd62 fix(ui): render dataset run description and metadata on screen given table (#4089)
fix(ui): render dataset run description and metadata on screen without cutting off table
2024-11-06 19:47:01 +00:00
Steffen SchmitzandGitHub 5d4125208c chore: restrict domain timestamp updates to one day in clickhouse ingestion (#4085) 2024-11-06 17:27:25 +00:00
Steffen SchmitzandGitHub ff3bfe75ec chore: drop redundant observation project_id index on Clickhouse (#4084) 2024-11-06 16:45:53 +01:00
Steffen SchmitzandGitHub 89c31fdbc2 chore: add type filter to observation lookup in ingestion queue (#4083) 2024-11-06 15:18:19 +00:00
Steffen SchmitzandGitHub 9445c7ca78 chore: forward /api/public/scores to /api/public/ingestion (#4064) 2024-11-06 11:32:38 +01:00
marliessophieandGitHub b3d5833520 fix: unify ordering of dataset_items and dataset_run_items (#4063)
* fix: unify ordering of `dataset_items` and `dataset_run_items`

* fix: filter trace scores by observation id NULL
2024-11-05 19:03:30 +00:00
Steffen SchmitzandGitHub 20ed40510a chore: update cost calculation to handle null values from DB (#4057) 2024-11-05 18:22:10 +01:00
Steffen SchmitzandGitHub 9c5c57f134 chore: add test case for different input/output values (#4061) 2024-11-05 16:46:17 +00:00
Steffen SchmitzandGitHub 09b916aac9 chore: do not send auth errors to DLQ in bull (#4062) 2024-11-05 16:23:17 +00:00
marliessophieandGitHub 92843fef28 feat(datasets): view to compare different dataset runs (#3971) 2024-11-05 16:10:28 +01:00
Steffen SchmitzandGitHub 433951afda chore: retry legacy ingestion batches on errors (#4058) 2024-11-05 15:25:25 +01:00
Steffen SchmitzandGitHub aa1f861e4e chore: ignore all .env files aside from examples (#4059) 2024-11-05 13:25:27 +00:00
Hassieb PakzadandGitHub 66b6a09400 feat(models): add claude haiku 3.5 support (#4055) 2024-11-05 12:14:25 +01:00
Max DeichmannandGitHub edf0dc8399 feat: add metadata filter for all UI tables using clickhouse (#4022)
* feat: add metadata filter for all UI tables using clickhouse

* feat: add metadata filter for all UI tables using clickhouse

* push
2024-11-05 10:05:22 +00:00
Marc KlingenandGitHub 8d0c1219ca fix: env configuration for telemetry (#4052) 2024-11-04 23:04:07 +00:00
Max DeichmannandGitHub 33b3f69076 feat: query distinct models for charts (#4051) 2024-11-04 22:42:46 +01:00
Max DeichmannandGitHub f582941ad3 perf: fix timeseries clickouse (#4049) 2024-11-04 21:14:14 +00:00
Max DeichmannandGitHub 9fd57b4fbf feat: query second row of dashboards using clickhouse (#4048) 2024-11-04 20:18:37 +00:00
Max DeichmannandGitHub f61f69c238 perf: only join in dashboard if required (#4047) 2024-11-04 18:17:22 +00:00
Steffen SchmitzandGitHub 55c626d627 chore: reduce startTime missing logs to create events (#4044) 2024-11-04 14:52:09 +00:00
Max DeichmannandGitHub c6d3257916 feat: add first three dashboards in clickhouse (#4023) 2024-11-04 15:02:29 +01:00
Marc KlingenandGitHub 34f91e9872 chore(ui): slightly lighter bg of darkmode sidebar (#4042) 2024-11-04 15:01:58 +01:00
Steffen SchmitzandGitHub 5acd509fc0 chore: init migration scripts with state on restarts (#4039) 2024-11-04 13:37:28 +00:00
marliessophieandGitHub 765756b002 fix: scores tab in observation preview to only show scores liked to observation and trace id (#4030) 2024-11-04 12:54:21 +00:00
Steffen SchmitzandGitHub 64525744de chore: include usagedetails from postgres in ingestion pipeline (#4034) 2024-11-04 12:41:47 +00:00
Steffen SchmitzandGitHub 2b89f62add chore: remove temporary column from PG to CH migrations (#4031) 2024-11-04 13:26:42 +01:00
marliessophieandGitHub 0a67adf3e7 fix: show generations count on prompt metrics table (#4003) 2024-11-04 09:56:18 +00:00
marliessophieandGitHub 4ef7296e37 fix: multi-select header partially cut off on Firefox (#4029) 2024-11-04 09:45:18 +00:00
Max Deichmann 0cef46cd78 chore: release v2.87.0
CI/CD / lint (push) Waiting to run
CI/CD / test-docker-build (push) Waiting to run
CI/CD / tests-web (node20, pg12) (push) Waiting to run
CI/CD / tests-web (node20, pg15) (push) Waiting to run
CI/CD / tests-worker (node20, pg12) (push) Waiting to run
CI/CD / tests-worker (node20, pg15) (push) Waiting to run
CI/CD / e2e-tests (push) Waiting to run
CI/CD / e2e-server-tests (push) Waiting to run
CI/CD / all-ci-passed (push) Blocked by required conditions
CI/CD / push-docker-image (push) Blocked by required conditions
release.yml / release (push) Waiting to run
2024-11-03 18:34:51 +01:00
Max DeichmannandGitHub 243d72772c perf: improve select by id observations performance (#4021) 2024-11-03 16:45:43 +00:00
Max DeichmannandGitHub 91214cc6a6 feat: add order by to clickhouse tables (#4020) 2024-11-03 16:35:05 +00:00
Max DeichmannandGitHub da85e80689 feat: add generations table clickhouse support (#4015) 2024-11-03 14:48:10 +01:00
Max DeichmannandGitHub e22827d380 perf: improve traces metrics API performance (#4017) 2024-11-02 21:03:51 +00:00
Marc KlingenandGitHub 7161643574 docs: add star history to readme 2024-11-02 21:52:31 +01:00
Max DeichmannandGitHub 77f1c6cd0c fix: respect limit and offset for traces table query (#4016) 2024-11-02 19:38:06 +01:00
Max DeichmannandGitHub 0a4b56eb28 chore: improve clickhouse observability (#4014) 2024-11-02 17:50:26 +00:00
Max DeichmannandGitHub b5f94d78f3 fix: fix clickhouse untis and conversions (#4013) 2024-11-02 17:39:52 +00:00
Max DeichmannandGitHub a9fd1e1832 fix: fix timestamps in clickhouse reading and refactor clickhouse reads (#4011) 2024-11-02 15:53:37 +01:00
Max DeichmannandGitHub 762a7e3489 chore: fix fetching model for clickhouse reads (#4008) 2024-11-02 11:02:13 +00:00
Max DeichmannandGitHub 0f3717db0d fix: fix eval execution retries (#4010) 2024-11-02 10:22:19 +00:00
Max DeichmannandGitHub 8590ab84ba fix: fix traces table latency (#4005) 2024-11-01 16:49:48 +01:00
marliessophieandGitHub 9d17903392 fix: select all functionality on prompt metrics table (#4004) 2024-11-01 15:16:24 +00:00
Max DeichmannandGitHub 7d94831c01 feat: support for traces detail view with clickhouse (#3982) 2024-11-01 14:51:59 +00:00
Marc KlingenandGitHub b97900f5b8 perf(ui): do not refetch filter options on page mount (#4002) 2024-11-01 14:38:58 +00:00
marliessophieandGitHub 887947a967 fix: opening public langfuse urls (#3999)
* fix: opening public langfuse urls

* push
2024-11-01 14:14:07 +00:00
Max DeichmannandGitHub 85661c0e58 fix: remove final from ingestion (#3998) 2024-11-01 13:21:46 +00:00
fc2ba97792 feat(ui): add new sidebar and main navigation (#3837)
Co-authored-by: Marlies Mayerhofer <74332854+marliessophie@users.noreply.github.com>
2024-11-01 12:12:47 +01:00
Max DeichmannandGitHub 4b6cd3fd2b revert: revert clickhouse trace check (#3993) 2024-10-31 23:54:08 +01:00
Max DeichmannandGitHub d758821817 feat: support traces.countAll api with clickhouse (#3974) 2024-10-31 20:35:18 +00:00
Steffen SchmitzandGitHub 4685d0b1b6 chore: ensure model parameters are cast correctly in clickhouse ingestion (#3992) 2024-10-31 20:14:13 +00:00
Max DeichmannandGitHub e88946f0b5 fix: fix zod schema (#3991) 2024-10-31 19:45:56 +01:00
Steffen SchmitzandGitHub f14c611aac chore: correct number parsing from postgres in clickhouse merge (#3987) 2024-10-31 17:07:58 +00:00
Max DeichmannandGitHub 84ccdc24ed feat: log all project_ids with missing domain id (#3984) 2024-10-31 16:29:39 +00:00
Steffen SchmitzandGitHub 04a0366af5 chore: add fallback to event timestamp again, log warning (#3983) 2024-10-31 15:43:19 +00:00
Max DeichmannandGitHub 1a956c909c feat: add additional traces table filter (#3973) 2024-10-31 16:28:09 +01:00
Steffen SchmitzandGitHub 8e150701cb chore: join postgres state into Clickhouse results and add event_ts (#3978) 2024-10-31 15:44:26 +01:00
Max DeichmannandGitHub b9d8b5026e feat: add clickhouse support for traces filter options endpoint (#3969) 2024-10-30 21:05:31 +00:00
Max DeichmannandGitHub a6dfc58c72 feat: read traces and observations from clickhouse (#3810) 2024-10-30 21:32:08 +01:00
Steffen SchmitzandGitHub 6266cd8f0b chore: update traces and observations tables to use map type, full mapping script for migration (#3960) 2024-10-30 19:51:24 +00:00
Steffen SchmitzandGitHub e6e6328594 chore: remove remaining zod parses in worker pipeline (#3968) 2024-10-30 11:05:24 +00:00
Steffen SchmitzandGitHub 8b8826437a chore: remove zod parsing from worker (#3967) 2024-10-30 09:21:07 +00:00
marliessophieandGitHub b77c51512e feat(evals): option to update ref eval configs when creating new version (#3919) 2024-10-30 09:54:44 +01:00
marliessophieandGitHub 10d2299fae chore: global error handling for trpc routes (#3959)
* refactor: extract middleware for trpc routers to consistently throw trpc error

* chore: drop user-facing message in case of trpc error to not expose internals

* drop cause from error object
2024-10-29 19:02:59 +01:00
marliessophieandGitHub 83e1eb50bd Revert "chore(deps): bump react-hook-form from 7.51.5 to 7.53.0 (#3917)" (#3958)
Revert "chore(deps): bump react-hook-form from 7.51.5 to 7.53.0 (#3613)"
2024-10-29 14:24:26 +00:00
Hassieb PakzadandGitHub c7546bcafd fix(models): upsert prices on model drift; (#3955) 2024-10-29 14:39:10 +01:00
b26a537181 refactor: full-screen-page to drop magic in height calculation (#3944)
* fix: remove magic from full screen page

* fix: wrap prompts table as full page

* style: account for sticky header layout for mobile
---------

Co-authored-by: Marc Klingen <git@marcklingen.com>
2024-10-29 10:34:45 +00:00
marliessophieandGitHub e5cbca7f95 style(evals): add table metadata view for eval config detail (#3916) 2024-10-29 09:18:54 +00:00
Max DeichmannandGitHub 73c7a554c2 chore: add external packages to next config (#3950) 2024-10-29 10:07:14 +01:00
Steffen SchmitzandGitHub c460029704 chore: add additional requirements for background migrations (#3951) 2024-10-29 08:45:20 +00:00
Max DeichmannandGitHub 2808157200 chore: add bullmq to external package (#3947) 2024-10-28 22:21:11 +00:00
Max DeichmannandGitHub 79913e3816 chore: adjust eval execution retries (#3946) 2024-10-28 22:58:31 +01:00
Marc KlingenandGitHub c9b79f4277 chore: pnpm dx pulls latest infra containers, waits for healthcheck, prunes old volumes (#3945) 2024-10-28 20:03:51 +01:00
Steffen SchmitzandGitHub 5897db3880 chore: add background migrations for long-running tasks (#3895) 2024-10-28 14:24:08 +00:00
Steffen SchmitzandGitHub 4a4d08d9f1 chore: create v3beta docker compose file (#3936) 2024-10-28 14:10:33 +00:00
marliessophieandGitHub f383a3f50f chore: support table-level-tab navigation through eval configs, templates, log (#3899) 2024-10-28 14:55:51 +01:00
Max DeichmannandGitHub 79954cf188 fix: treat api call failures for eval executions the same way as parsing errors. (#3942) 2024-10-28 12:59:54 +00:00
Max DeichmannandGitHub f87adc209a fix: do not fail eval queue on LLM Api errors (#3939) 2024-10-28 11:50:02 +00:00
marliessophieandGitHub 3a71bfd763 fix: populate default metadata when adding observation as dataset item (#3938) 2024-10-28 11:22:53 +00:00
Steffen SchmitzandGitHub c8ff61057e chore: remove hasCreateEvent check in ingestion pipeline (#3937) 2024-10-28 10:01:28 +00:00
Steffen SchmitzandGitHub 10d5975e26 chore: add is_deleted column to clickhouse seeder (#3934) 2024-10-28 08:14:03 +00:00
Steffen SchmitzandGitHub 7b34c6fc32 chore(deps): dependency updates (#3917) 2024-10-25 15:34:22 +02:00
Steffen SchmitzandGitHub 1ac6e51271 chore: validate clickhouse data quality (#3809) 2024-10-25 12:31:45 +00:00
Steffen SchmitzandGitHub 8435080ec7 chore: remove prod-eu-temp deploy environment (#3849) 2024-10-25 09:22:06 +00:00
Marc KlingenandGitHub 3b474e596d chore: faster healthcheck of minio in dev docker compose (#3914) 2024-10-25 11:11:45 +02:00
marliessophieandGitHub 16b40917c6 feat(evals): allow creating template from eval config form (#3896) 2024-10-25 08:39:00 +00:00
Hassieb Pakzad d0e5a6bfe3 chore(models): increase upsert diff tolerance 2024-10-24 19:48:52 +02:00
Hassieb Pakzad 0048dbfae2 docs(models): add note on how to update default models 2024-10-24 19:01:25 +02:00
Hassieb PakzadandGitHub bc612a8384 fix(models): allow for time tolerance in updateat diff (#3903) 2024-10-24 18:47:11 +02:00
Hassieb PakzadandGitHub 541765f8a0 feat(cost): make cost tracking flexible (#3811) 2024-10-24 18:12:38 +02:00
Marc KlingenandGitHub 7b4663e373 feat(cloud): add event usage metering (#3898)
push
2024-10-24 15:25:37 +02:00
marliessophieandGitHub 5eb792e2f5 fix(evals): show eval filters in eval-configs table (#3897) 2024-10-24 13:03:49 +00:00
Max DeichmannandGitHub 5499da8cc3 feat: add is_deleted to clickhouse (#3887) 2024-10-24 10:01:49 +00:00
Hassieb PakzadandGitHub 35b212954c fix(models): update sonnet model name (#3894) 2024-10-24 11:48:34 +02:00
Max DeichmannandGitHub 11c791ae38 fix: improve ingestion pipeline by id query (#3883) 2024-10-23 20:50:53 +02:00
Max DeichmannandGitHub 4ab1cbd03c perf: remove query cache in ingestion pipeline (#3882) 2024-10-23 15:48:34 +00:00
Max DeichmannandGitHub d33fb015d5 feat: adjust clickhouse writer (#3881) 2024-10-23 16:45:18 +02:00
Max DeichmannandGitHub 5b1704c95e perf: improve CH primary keys and order by (#3880) 2024-10-23 14:26:50 +00:00
Max DeichmannandGitHub 03659f3e82 feat: remove io from wide tables (#3878) 2024-10-23 15:45:28 +02:00
Max DeichmannandGitHub cf3b7f4603 chore: add clickhouse writer attributes (#3867) 2024-10-22 21:31:49 +00:00
Max Deichmann 07c6c95eff chore: release v2.86.0
CI/CD / lint (push) Waiting to run
CI/CD / test-docker-build (push) Waiting to run
CI/CD / tests-web (node20, pg12) (push) Waiting to run
CI/CD / tests-web (node20, pg15) (push) Waiting to run
CI/CD / tests-worker (node20, pg12) (push) Waiting to run
CI/CD / tests-worker (node20, pg15) (push) Waiting to run
CI/CD / e2e-tests (push) Waiting to run
CI/CD / e2e-server-tests (push) Waiting to run
CI/CD / all-ci-passed (push) Blocked by required conditions
CI/CD / push-docker-image (push) Blocked by required conditions
release.yml / release (push) Waiting to run
2024-10-22 20:49:49 +02:00
Max DeichmannandGitHub 74cbe04fe9 feat: add new claude model (#3865) 2024-10-22 20:42:35 +02:00
marliessophieandGitHub 05aa62aa9e style: unify dataset buttons/badge with langfuse design system, add UI access check on archive dataset item (#3864)
* style: `New dataset` button as `secondary`

* fix: missing UI-only access check on archive item

* style button consistently
2024-10-22 18:30:27 +00:00
marliessophieandGitHub 4acb01086b fix: don't throw unauthorized errors when viewing a public trace (#3852) 2024-10-22 11:59:54 +02:00
Max DeichmannandGitHub d12941576b feat: format logs (#3860) 2024-10-22 10:48:56 +02:00
Max DeichmannandGitHub d99f12fe64 feat: improve clickhouse seeder (#3853) 2024-10-22 09:02:01 +02:00
marliessophieandGitHub 187e64c8d2 fix(ui): table pagination distinguish between no data and loading states (#3851) 2024-10-21 17:33:37 +02:00
Max DeichmannandGitHub de152e4728 chore: remove prettier config (#3806)
fix
2024-10-21 14:20:48 +00:00
Steffen SchmitzandGitHub d3669eeef2 chore: add prod-eu deployment option (#3847) 2024-10-21 16:07:02 +02:00
Max DeichmannandGitHub 38a453eebb fix: local postgres query logging (#3846) 2024-10-21 13:04:44 +00:00
marliessophieandGitHub 53dc3c3473 feat(api): add traceTags query param to GET /scores (#3843)
* feat(api): add `traceTags` query param to GET /scores

* docs
2024-10-21 14:11:34 +02:00
Max DeichmannandGitHub b573d9035f chore: add tokenisation metrics (#3830) 2024-10-18 15:21:22 -07:00
Steffen SchmitzandGitHub bc0485c91a chore: publish custom metrics to CloudWatch (#3828) 2024-10-18 14:45:08 -07:00
Marc KlingenandGitHub 63fa223e54 bug: UI prompt edit changes go away when tab switching (#3827)
Fixes langfuse/langfuse#3807
2024-10-18 18:29:03 +00:00
marliessophieandGitHub 06b3839fd2 feat: support DnD reordering of playground messages (#3826) 2024-10-18 18:14:26 +00:00
Marc KlingenandGitHub 254c399ec6 fix: use trace/observation createdAt for health endpoint (#3825) 2024-10-18 16:38:08 +00:00
Max DeichmannandGitHub 6567e93a96 chore: add traces wide table (#3808) 2024-10-17 22:42:34 +00:00
Max DeichmannandGitHub 7f6de12256 chore: set ingestion delay to zero (#3804) 2024-10-17 20:15:44 +00:00
Max DeichmannandGitHub 5090e521a1 fix: only error eval execution for unexpected cases (#3805) 2024-10-17 19:12:54 +00:00
marliessophieandGitHub 81847e7bc5 feat(ui): support comments on annotation queue view (#3792)
* feat: support comments on annotation queue view

* refactor: queue header accessibility
2024-10-17 18:59:54 +00:00
marliessophieandGitHub b4fff64169 feat(ui): render plain string as markdown by default on single trace/observation page (#3788)
* feat: render plain string as markdown by default

* fix: markdown button margins
2024-10-17 16:45:48 +00:00
Max DeichmannandGitHub 3cba71d5be feat: add ingestion delay (#3803) 2024-10-17 16:29:08 +00:00
marliessophieandGitHub ce6a96f3d1 fix(tables): serialization issue in score column visibility (#3793)
* previously score columns including a "." in name could not be hidden
2024-10-17 01:38:54 +00:00
Steffen SchmitzandGitHub f7a00ff0eb chore: add error listeners on queues (#3791) 2024-10-16 17:54:28 -07:00
Steffen SchmitzandGitHub 56306bb456 chore: reconnect redis on errors (#3785) 2024-10-16 23:29:08 +00:00
Steffen SchmitzandGitHub eab3268386 chore: fix queue metric names (#3790) 2024-10-16 22:37:04 +00:00
marliessophieandGitHub f68dfac8b4 fix: don't use accessible title for queue item (#3787) 2024-10-16 22:08:37 +00:00
Steffen SchmitzandGitHub bb76ae4aea chore(refactor): calculate queue metrics in WorkerManager (#3732) 2024-10-16 21:55:01 +00:00
marliessophieandGitHub 7babcbd170 feat(api): POST /comments GET /comments GET /comments/${commentId} (#3383) 2024-10-16 23:22:23 +02:00
Hassieb PakzadandGitHub ef12dd1b5f fix(evals): put reasoning key before score (#3784) 2024-10-16 12:00:49 -07:00
Max DeichmannandGitHub 6200a9844f chore: improve clickhouse schema (#3638) 2024-10-16 18:00:45 +00:00
Steffen SchmitzandGitHub 1e5717d13d chore: batch events by eventBodyId for reduces S3 interactions (#3770) 2024-10-16 17:48:24 +00:00
Hassieb PakzadandGitHub c97ed69874 chore(prompts): config data type to JSON (#3759) 2024-10-16 10:36:01 -07:00
marliessophieandGitHub c258776a68 feat(api): add queueId query param to GET /scores (#3771) 2024-10-16 17:29:13 +00:00
marliessophieandGitHub 892ff766e8 fix(ui): transition from uncontrolled to controlled component, add screen reader drawer title (#3774) 2024-10-16 19:16:59 +02:00
marliessophieandGitHub 80470cb812 chore: support consistent empty and loading states in dashboard (#3763) 2024-10-16 05:26:27 +02:00
Steffen SchmitzandGitHub 5949a659a9 chore: remove trace sampler on worker (#3762) 2024-10-15 21:00:02 +00:00
marliessophieandGitHub 151d305b02 chore: support more decimal digits for cost in session view (#3758) 2024-10-15 16:37:17 +00:00
Marc KlingenandGitHub 74819a1d1f fix(self-host): LANGFUSE_INIT_USER_EMAIL lowercased for headless initialization (#3754) 2024-10-15 07:01:50 +00:00
marliessophieandGitHub 45edcde550 style(ui): standardize table header line height across selection states (#3753) 2024-10-15 03:03:06 +00:00
marliessophieandGitHub 2b04b9d9b8 fix(user_table): reset orderBy state on tab navigation (#3751) 2024-10-15 01:35:33 +02:00
marliessophieandGitHub 3150c655fd fix: loading state of annotation queue dropdown (#3750) 2024-10-14 23:04:11 +00:00
Hassieb PakzadandGitHub 26375885a9 fix(llm-api-keys): add exact bedrock permission needed (#3749) 2024-10-14 15:43:37 -07:00
marliessophieandGitHub b584684593 fix: add default minimum width for dynamic score columns, allow column reordering for prompts metrics table (#3747)
* fix: define fallback width for dynamic columns

* feat: allow column reordering on metrics table
2024-10-14 22:13:32 +00:00
marliessophieandGitHub 6f00b67055 feat: add select all / deselect all button to all tables columns visibility picker (#3740)
* feat: add select all / deselect all button to all tables columns picker

* fix: don't hide columns that we shouldn't
2024-10-14 23:53:06 +02:00
Max Deichmann cdbae876e8 chore: release v2.85.1
CI/CD / lint (push) Waiting to run
CI/CD / test-docker-build (push) Waiting to run
CI/CD / tests-web (node20, pg12) (push) Waiting to run
CI/CD / tests-web (node20, pg15) (push) Waiting to run
CI/CD / tests-worker (node20, pg12) (push) Waiting to run
CI/CD / tests-worker (node20, pg15) (push) Waiting to run
CI/CD / e2e-tests (push) Waiting to run
CI/CD / e2e-server-tests (push) Waiting to run
CI/CD / all-ci-passed (push) Blocked by required conditions
CI/CD / push-docker-image (push) Blocked by required conditions
release.yml / release (push) Waiting to run
2024-10-14 12:34:33 -07:00
Max DeichmannandGitHub d1c61b4eba chore: upgrade json path (#3741) 2024-10-14 19:32:27 +00:00
Steffen SchmitzandGitHub e443730e66 chore: group events on S3 by clickhouse entity (#3733) 2024-10-14 18:34:06 +00:00
marliessophieandGitHub 4c55165c2f fix: show number of observations in trace table (#3743) 2024-10-14 18:22:12 +00:00
Hassieb PakzadandGitHub 34a4a22cba fix(models): start_date type to datetime in fern (#3742) 2024-10-14 10:49:01 -07:00
Steffen SchmitzandGitHub 4cb688929e chore: add otel trace sampling support (#3739) 2024-10-14 17:25:37 +00:00
Steffen SchmitzandGitHub 0b86f3207b chore: authenticate with docker hub for CI (#3734) 2024-10-14 09:45:20 -07:00
Steffen SchmitzandGitHub 4a83749425 chore: migrate clickhouse ingestion to S3 based events (#3719) 2024-10-14 14:48:26 +00:00
356 changed files with 21837 additions and 8367 deletions
+3 -4
View File
@@ -37,15 +37,15 @@ SMTP_CONNECTION_URL="" # Defines the connection url for smtp server.
S3_ENDPOINT=http://localhost:9090
S3_ACCESS_KEY_ID=minio
S3_SECRET_ACCESS_KEY=miniosecret
S3_BUCKET_NAME=mybucket
S3_BUCKET_NAME=langfuse
S3_REGION=us-east-1
## Necessary for minio compatibility
S3_FORCE_PATH_STYLE=true
# S3 Event Bucket Upload
## Set to true to test uploading all events to S3
LANGFUSE_S3_EVENT_UPLOAD_ENABLED=false
LANGFUSE_S3_EVENT_UPLOAD_BUCKET=mybucket
LANGFUSE_S3_EVENT_UPLOAD_ENABLED=true
LANGFUSE_S3_EVENT_UPLOAD_BUCKET=langfuse
LANGFUSE_S3_EVENT_UPLOAD_ACCESS_KEY_ID=minio
LANGFUSE_S3_EVENT_UPLOAD_SECRET_ACCESS_KEY=miniosecret
LANGFUSE_S3_EVENT_UPLOAD_REGION=us-east-1
@@ -62,7 +62,6 @@ REDIS_HOST="127.0.0.1"
REDIS_PORT=6379
REDIS_AUTH="myredissecret"
LANGFUSE_WORKER_PASSWORD=mybasicauthsecret
# openssl rand -hex 32 used only here
ENCRYPTION_KEY=0000000000000000000000000000000000000000000000000000000000000000
-2
View File
@@ -18,7 +18,5 @@ REDIS_HOST="127.0.0.1"
REDIS_PORT=6379
REDIS_AUTH="myredissecret"
LANGFUSE_WORKER_PASSWORD=myworkerpassword
# openssl rand -hex 32 used only here
ENCRYPTION_KEY=0000000000000000000000000000000000000000000000000000000000000000
+2 -5
View File
@@ -237,11 +237,8 @@ OTEL_SERVICE_NAME="langfuse"
# CLICKHOUSE_PASSWORD=
# Ingestion
# LANGFUSE_INGESTION_BUFFER_TTL_SECONDS=
# LANGFUSE_INGESTION_FLUSH_DELAY_MS=
# LANGFUSE_INGESTION_FLUSH_ATTEMPTS=
# LANGFUSE_INGESTION_FLUSH_PROCESSING_CONCURRENCY=
# LANGFUSE_INGESTION_CLICKHOUSE_WRITE_BATCH_SIZE=
# LANGFUSE_INGESTION_QUEUE_DELAY_MS=
# LANGFUSE_INGESTION_CLICKHOUSE_WRITE_BATCH_SIZE=
# LANGFUSE_INGESTION_CLICKHOUSE_WRITE_INTERVAL_MS=
# LANGFUSE_INGESTION_CLICKHOUSE_MAX_ATTEMPTS=
# LANGFUSE_LEGACY_INGESTION_WORKER_CONCURRENCY=
+2 -2
View File
@@ -19,7 +19,7 @@ on:
type: choice
options:
- staging
- prod-eu-temp
- prod-eu
- prod-us
required: true
@@ -78,7 +78,7 @@ jobs:
return `["staging"]`
}
if (context.ref === "refs/heads/production") {
return `["prod-eu-temp", "prod-us"]`
return `["prod-eu", "prod-us"]`
}
}
return "[]"
+32 -47
View File
@@ -40,15 +40,17 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Login to Docker Hub
uses: docker/login-action@v3
with:
username: ${{ secrets.DOCKERHUB_USERNAME_READ }}
password: ${{ secrets.DOCKERHUB_TOKEN_READ }}
- name: Set NEXT_PUBLIC_BUILD_ID
run: echo "NEXT_PUBLIC_BUILD_ID=$(git rev-parse --short HEAD)" >> $GITHUB_ENV
- name: Build and run both images from compose
run: |
docker compose -f docker-compose.build.yml up -d
sleep 5 # Wait for PostgreSQL to accept connections
- name: Ensure no unhealthy status
run: |
if docker-compose ps | grep "(unhealthy)"; then
@@ -57,11 +59,9 @@ jobs:
else
echo "All services are healthy"
fi
- name: Check worker health
run: |
timeout 10 bash -c 'until curl -f http://localhost:3030/api/health; do sleep 2; done'
- name: Check server health
run: |
timeout 10 bash -c 'until curl -f http://localhost:3000/api/public/health; do sleep 2; done'
@@ -79,27 +79,33 @@ jobs:
uses: pierotofy/set-swap-space@master
with:
swap-size-gb: 10
- uses: actions/checkout@v4
- name: Install golang-migrate for Clickhouse migrations
run: |
curl -L https://github.com/golang-migrate/migrate/releases/download/v4.16.2/migrate.linux-amd64.tar.gz | tar xvz
sudo mv migrate /usr/bin/migrate
which migrate
- uses: pnpm/action-setup@v3
with:
version: 9.5.0
- name: Login to Docker Hub
uses: docker/login-action@v3
with:
username: ${{ secrets.DOCKERHUB_USERNAME_READ }}
password: ${{ secrets.DOCKERHUB_TOKEN_READ }}
- name: Use Node.js ${{ matrix.node-version }}
uses: actions/setup-node@v4
with:
node-version: ${{ matrix.node-version }}
cache: "pnpm"
cache-dependency-path: "pnpm-lock.yaml"
- name: install dependencies
run: |
pnpm install
- name: Load default env
run: |
cp .env.dev.example .env
grep -v -e '^S3_BUCKET_NAME=' -e '^REDIS_HOST=' -e '^NEXT_PUBLIC_LANGFUSE_RUN_NEXT_INIT=' .env.dev.example > .env
- name: Run + migrate
run: |
docker compose -f docker-compose.dev.yml up -d
@@ -107,14 +113,12 @@ jobs:
docker compose ps
env:
POSTGRES_VERSION: ${{ matrix.postgres-version }}
- name: Seed DB
run: |
pnpm run db:migrate
pnpm --filter=shared ch:up
- name: Build
run: pnpm run build
- name: Start Langfuse
run: (pnpm run start&)
env:
@@ -127,7 +131,6 @@ jobs:
LANGFUSE_INIT_USER_EMAIL: "demo@langfuse.com"
LANGFUSE_INIT_USER_NAME: "Demo User"
LANGFUSE_INIT_USER_PASSWORD: "password"
- name: run tests
run: pnpm --filter=web run test
@@ -144,41 +147,39 @@ jobs:
uses: pierotofy/set-swap-space@master
with:
swap-size-gb: 10
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v3
with:
version: 9.5.0
- name: Login to Docker Hub
uses: docker/login-action@v3
with:
username: ${{ secrets.DOCKERHUB_USERNAME_READ }}
password: ${{ secrets.DOCKERHUB_TOKEN_READ }}
- name: Use Node.js ${{ matrix.node-version }}
uses: actions/setup-node@v4
with:
node-version: ${{ matrix.node-version }}
cache: "pnpm"
cache-dependency-path: "pnpm-lock.yaml"
- name: install dependencies
run: |
pnpm install
- name: Install golang-migrate for Clickhouse migrations
run: |
curl -L https://github.com/golang-migrate/migrate/releases/download/v4.16.2/migrate.linux-amd64.tar.gz | tar xvz
sudo mv migrate /usr/bin/migrate
which migrate
- name: Load default env
run: |
cp .env.dev.example .env
cp .env.dev.example web/.env
cp .env.dev.example worker/.env
- name: Run + migrate
run: |
docker compose -f docker-compose.dev.yml up -d
sleep 5 # Wait for PostgreSQL to accept connections
docker compose ps
- name: Ensure no unhealthy status
run: |
if docker compose ps | grep "(unhealthy)"; then
@@ -187,18 +188,16 @@ jobs:
else
echo "All services are healthy"
fi
- name: Seed DB
run: |
pnpm run db:migrate
pnpm run db:seed
pnpm run --filter=shared ch:up
- name: Build
run: pnpm --filter=worker... run build
- name: run tests
run: pnpm --filter=worker run test
e2e-tests:
runs-on: ubuntu-latest
steps:
@@ -211,33 +210,31 @@ jobs:
node-version: 20
cache: "pnpm"
cache-dependency-path: "pnpm-lock.yaml"
- name: Login to Docker Hub
uses: docker/login-action@v3
with:
username: ${{ secrets.DOCKERHUB_USERNAME_READ }}
password: ${{ secrets.DOCKERHUB_TOKEN_READ }}
- name: install dependencies
run: |
pnpm install
- name: Load default env
run: |
cp .env.dev.example .env
cp .env.dev.example web/.env
- name: Run + migrate
run: |
docker compose -f docker-compose.dev.yml up -d
docker compose ps
sleep 5 # Wait for PostgreSQL to accept connections
- name: Seed DB
run: |
pnpm run db:migrate
pnpm run db:seed
- name: Build
run: pnpm run build
- name: Install playwright
run: pnpm --filter=web exec playwright install --with-deps
- name: Run e2e tests
run: pnpm --filter=web run test:e2e
@@ -245,6 +242,11 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Login to Docker Hub
uses: docker/login-action@v3
with:
username: ${{ secrets.DOCKERHUB_USERNAME_READ }}
password: ${{ secrets.DOCKERHUB_TOKEN_READ }}
- uses: pnpm/action-setup@v3
with:
version: 9.5.0
@@ -253,35 +255,28 @@ jobs:
node-version: 20
cache: "pnpm"
cache-dependency-path: "pnpm-lock.yaml"
- name: install dependencies
run: |
pnpm install
- name: Load default env
run: |
cp .env.dev.example .env
echo "LANGFUSE_ASYNC_INGESTION_PROCESSING=true" >> .env
echo "LANGFUSE_CACHE_API_KEY_ENABLED=true" >> .env
echo "LANGFUSE_CACHE_PROMPT_ENABLED=true" >> .env
- name: Run + migrate
run: |
docker compose -f docker-compose.dev.yml up -d
docker compose ps
sleep 5 # Wait for PostgreSQL to accept connections
- name: Seed DB
run: |
pnpm run db:migrate
pnpm run db:seed:examples
- name: Build
run: pnpm run build
- name: Run server
run: (pnpm run start&)
- name: Run e2e tests
run: pnpm --filter=web run test:e2e:server
@@ -326,34 +321,27 @@ jobs:
with:
node-version: 20
cache-dependency-path: "pnpm-lock.yaml"
- name: Checkout
uses: actions/checkout@v4
- name: Set NEXT_PUBLIC_BUILD_ID
run: echo "NEXT_PUBLIC_BUILD_ID=$(git rev-parse --short HEAD)" >> $GITHUB_ENV
- name: Log in to the GitHub Container registry
uses: docker/login-action@v2
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Log in to Docker Hub
uses: docker/login-action@v2
with:
username: ${{ secrets.DOCKERHUB_USERNAME }}
password: ${{ secrets.DOCKERHUB_TOKEN }}
- name: Set up QEMU
uses: docker/setup-qemu-action@v3
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
with:
driver-opts: network=host
- name: Extract metadata (tags, labels) for Docker
id: meta-web
uses: docker/metadata-action@v4
@@ -368,7 +356,6 @@ jobs:
type=semver,pattern={{version}}
type=semver,pattern={{major}}.{{minor}}
type=semver,pattern={{major}}
- name: Build and push Docker image (web)
uses: docker/build-push-action@v4
with:
@@ -380,7 +367,6 @@ jobs:
platforms: |
linux/amd64
${{ startsWith(github.ref, 'refs/tags/') && 'linux/arm64' || '' }}
- name: Extract metadata (tags, labels) for Docker
id: meta-worker
uses: docker/metadata-action@v4
@@ -395,7 +381,6 @@ jobs:
type=semver,pattern={{version}}
type=semver,pattern={{major}}.{{minor}}
type=semver,pattern={{major}}
- name: Build and push Docker image (worker)
uses: docker/build-push-action@v4
with:
+4 -3
View File
@@ -35,9 +35,10 @@ yarn-error.log*
# local env files
# do not commit any .env files to git, except for the .env.example file. https://create.t3.gg/en/usage/env-variables#using-environment-variables
.env
.env.local
.env*.local
.env*
!.env.dev.example
!.env.local.example
!.env.prod.example
# vercel
.vercel
+1
View File
@@ -20,6 +20,7 @@
"prettier.documentSelectors": [
"**/*.{cjs,mjs,ts,tsx,astro,md,mdx,json,yaml,yml}"
],
"prettier.trailingComma": "all",
"mdx.experimentalLanguageServer": true,
"typescript.preferences.importModuleSpecifier": "non-relative",
"docwriter.style": "JSDoc",
+14
View File
@@ -432,6 +432,20 @@ Example:
op run --env-file="./.env" -- pnpm --filter=shared run db:deploy
```
### Editing default models and prices
You can update the default AI models and prices by adding or updating an entry in `worker/src/constants/default-model-prices.json`.
Please note that
- prices are in USD
- the list is ordered by ID, so make sure to keep this order
- the `updated_at` field must be updated with the current date in ISO 8601 format. Otherwise, the change will be ignored.
### Transition period until V3 release
Until the V3 release, both the JSON record must be updated **and** a migration must be created to continue supporting self-hosted users. Note that the migration must updated both the `models` as well as the `prices` table accordingly.
## License
Langfuse is MIT licensed, except for `ee/` folder. See [LICENSE](LICENSE) and [docs](https://langfuse.com/docs/open-source) for more details.
+10
View File
@@ -181,3 +181,13 @@ This helps us to:
None of the data is shared with third parties and does not include any sensitive information. We want to be super transparent about this and you can find the exact data we collect [here](/web/src/features/telemetry/index.ts).
You can opt-out by setting `TELEMETRY_ENABLED=false`.
### Star History
<a href="https://star-history.com/#langfuse/langfuse&Date">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/svg?repos=langfuse/langfuse&type=Date&theme=dark" />
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/svg?repos=langfuse/langfuse&type=Date" />
<img alt="Star History Chart" src="https://api.star-history.com/svg?repos=langfuse/langfuse&type=Date" />
</picture>
</a>
-3
View File
@@ -20,8 +20,6 @@ services:
- NEXTAUTH_URL=http://localhost:3000
- TELEMETRY_ENABLED=${TELEMETRY_ENABLED:-true}
- LANGFUSE_ENABLE_EXPERIMENTAL_FEATURES=${LANGFUSE_ENABLE_EXPERIMENTAL_FEATURES:-false}
- LANGFUSE_WORKER_HOST=${LANGFUSE_WORKER_HOST:-worker}
- LANGFUSE_WORKER_PASSWORD=${LANGFUSE_WORKER_PASSWORD:-mybasicauthsecret}
- LANGFUSE_INIT_ORG_ID=${LANGFUSE_INIT_ORG_ID:-}
- LANGFUSE_INIT_ORG_NAME=${LANGFUSE_INIT_ORG_NAME:-}
- LANGFUSE_INIT_PROJECT_ID=${LANGFUSE_INIT_PROJECT_ID:-}
@@ -58,7 +56,6 @@ services:
- REDIS_HOST=${REDIS_HOST:-redis}
- REDIS_PORT=${REDIS_PORT:-6379}
- REDIS_AUTH=${REDIS_AUTH:-myredissecret}
- LANGFUSE_WORKER_PASSWORD=${LANGFUSE_WORKER_PASSWORD:-mybasicauthsecret}
restart: always
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:3030/api/health"]
+10 -16
View File
@@ -20,7 +20,9 @@ services:
minio:
image: minio/minio
container_name: minio
command: server /data --console-address ":9001"
entrypoint: sh
# create the 'langfuse' bucket before starting the service
command: -c 'mkdir -p /data/langfuse && minio server --address ":9000" --console-address ":9001" /data'
environment:
MINIO_ACCESS_KEY: minio
MINIO_SECRET_KEY: miniosecret
@@ -31,23 +33,10 @@ services:
- langfuse_minio_data:/data
healthcheck:
test: ["CMD", "mc", "ready", "local"]
interval: 10s
interval: 1s
timeout: 5s
retries: 5
start_period: 5s
miniocreatebucket:
image: minio/mc
container_name: miniocreatebucket
entrypoint: ["/bin/sh", "-c"]
command: >
"mc alias set minio http://minio:9000 minio miniosecret &&
mc rm -r --force minio/mybucket || true &&
mc mb minio/mybucket &&
mc policy set download minio/mybucket"
depends_on:
minio:
condition: service_healthy
start_period: 1s
redis:
image: redis:7.2.4
@@ -60,6 +49,11 @@ services:
postgres:
image: postgres:${POSTGRES_VERSION:-latest}
restart: always
healthcheck:
test: ["CMD-SHELL", "pg_isready -U postgres"]
interval: 3s
timeout: 3s
retries: 10
command: ["postgres", "-c", "log_statement=all"]
environment:
- POSTGRES_USER=postgres
+131
View File
@@ -0,0 +1,131 @@
services:
langfuse-worker:
image: langfuse/langfuse-worker:latest
depends_on: &langfuse-depends-on
postgres:
condition: service_healthy
minio:
condition: service_healthy
redis:
condition: service_healthy
ports:
- "3030:3030"
environment: &langfuse-worker-env
DATABASE_URL: postgresql://postgres:postgres@postgres:5432/postgres
SALT: "mysalt"
ENCRYPTION_KEY: "0000000000000000000000000000000000000000000000000000000000000000" # generate via `openssl rand -hex 32`
TELEMETRY_ENABLED: ${TELEMETRY_ENABLED:-true}
LANGFUSE_ENABLE_EXPERIMENTAL_FEATURES: ${LANGFUSE_ENABLE_EXPERIMENTAL_FEATURES:-true}
LANGFUSE_ASYNC_INGESTION_PROCESSING: ${LANGFUSE_ASYNC_INGESTION_PROCESSING:-true}
LANGFUSE_ASYNC_CLICKHOUSE_INGESTION_PROCESSING: ${LANGFUSE_ASYNC_CLICKHOUSE_INGESTION_PROCESSING:-true}
CLICKHOUSE_URL: ${CLICKHOUSE_URL:-http://clickhouse:8123}
CLICKHOUSE_USER: ${CLICKHOUSE_USER:-clickhouse}
CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD:-clickhouse}
LANGFUSE_S3_EVENT_UPLOAD_ENABLED: ${LANGFUSE_S3_EVENT_UPLOAD_ENABLED:-true}
LANGFUSE_S3_EVENT_UPLOAD_BUCKET: ${LANGFUSE_S3_EVENT_UPLOAD_BUCKET:-langfuse}
LANGFUSE_S3_EVENT_UPLOAD_REGION: ${LANGFUSE_S3_EVENT_UPLOAD_REGION:-us-east-1}
LANGFUSE_S3_EVENT_UPLOAD_ACCESS_KEY_ID: ${LANGFUSE_S3_EVENT_UPLOAD_ACCESS_KEY_ID:-minio}
LANGFUSE_S3_EVENT_UPLOAD_SECRET_ACCESS_KEY: ${LANGFUSE_S3_EVENT_UPLOAD_SECRET_ACCESS_KEY:-miniosecret}
LANGFUSE_S3_EVENT_UPLOAD_ENDPOINT: ${LANGFUSE_S3_EVENT_UPLOAD_ENDPOINT:-http://minio:9000}
LANGFUSE_S3_EVENT_UPLOAD_FORCE_PATH_STYLE: ${LANGFUSE_S3_EVENT_UPLOAD_FORCE_PATH_STYLE:-true}
REDIS_HOST: ${REDIS_HOST:-redis}
REDIS_PORT: ${REDIS_PORT:-6379}
REDIS_AUTH: ${REDIS_AUTH:-myredissecret}
langfuse-web:
image: langfuse/langfuse:latest
depends_on: *langfuse-depends-on
ports:
- "3000:3000"
environment:
<<: *langfuse-worker-env
NEXTAUTH_URL: http://localhost:3000
NEXTAUTH_SECRET: mysecret
LANGFUSE_INIT_ORG_ID: ${LANGFUSE_INIT_ORG_ID:-}
LANGFUSE_INIT_ORG_NAME: ${LANGFUSE_INIT_ORG_NAME:-}
LANGFUSE_INIT_PROJECT_ID: ${LANGFUSE_INIT_PROJECT_ID:-}
LANGFUSE_INIT_PROJECT_NAME: ${LANGFUSE_INIT_PROJECT_NAME:-}
LANGFUSE_INIT_PROJECT_PUBLIC_KEY: ${LANGFUSE_INIT_PROJECT_PUBLIC_KEY:-}
LANGFUSE_INIT_PROJECT_SECRET_KEY: ${LANGFUSE_INIT_PROJECT_SECRET_KEY:-}
LANGFUSE_INIT_USER_EMAIL: ${LANGFUSE_INIT_USER_EMAIL:-}
LANGFUSE_INIT_USER_NAME: ${LANGFUSE_INIT_USER_NAME:-}
LANGFUSE_INIT_USER_PASSWORD: ${LANGFUSE_INIT_USER_PASSWORD:-}
clickhouse:
image: clickhouse/clickhouse-server
user: "101:101"
container_name: clickhouse
hostname: clickhouse
environment:
CLICKHOUSE_DB: default
CLICKHOUSE_USER: clickhouse
CLICKHOUSE_PASSWORD: clickhouse
volumes:
- langfuse_clickhouse_data:/var/lib/clickhouse
- langfuse_clickhouse_logs:/var/log/clickhouse-server
ports:
- "8123:8123"
- "9000:9000"
depends_on:
- postgres
minio:
image: minio/minio
container_name: minio
entrypoint: sh
# create the 'langfuse' bucket before starting the service
command: -c 'mkdir -p /data/langfuse && minio server --address ":9000" --console-address ":9001" /data'
environment:
MINIO_ROOT_USER: minio
MINIO_ROOT_PASSWORD: miniosecret
ports:
- "9090:9000"
- "9091:9001"
volumes:
- langfuse_minio_data:/data
healthcheck:
test: ["CMD", "mc", "ready", "local"]
interval: 1s
timeout: 5s
retries: 5
start_period: 1s
redis:
image: redis:7
restart: always
command: >
--requirepass ${REDIS_AUTH:-myredissecret}
ports:
- 6379:6379
healthcheck:
test: [ 'CMD', 'redis-cli', 'ping' ]
interval: 3s
timeout: 10s
retries: 10
postgres:
image: postgres:${POSTGRES_VERSION:-latest}
restart: always
healthcheck:
test: [ "CMD-SHELL", "pg_isready -U postgres" ]
interval: 3s
timeout: 3s
retries: 10
environment:
POSTGRES_USER: postgres
POSTGRES_PASSWORD: postgres
POSTGRES_DB: postgres
ports:
- 5432:5432
volumes:
- langfuse_postgres_data:/var/lib/postgresql/data
volumes:
langfuse_postgres_data:
driver: local
langfuse_clickhouse_data:
driver: local
langfuse_clickhouse_logs:
driver: local
langfuse_minio_data:
driver: local
+5
View File
@@ -44,5 +44,10 @@
"ts-node": "^10.9.2",
"tsc-watch": "^6.2.0",
"typescript": "^5.4.5"
},
"pnpm": {
"overrides": {
"jsonpath-plus": "10.0.0"
}
}
}
-4
View File
@@ -1,4 +0,0 @@
{
"trailingComma": "es5",
"printWidth": 120
}
+73
View File
@@ -0,0 +1,73 @@
# yaml-language-server: $schema=https://raw.githubusercontent.com/fern-api/fern/main/fern.schema.json
imports:
pagination: ./utils/pagination.yml
commons: ./commons.yml
service:
auth: true
base-path: /api/public
endpoints:
create:
docs: Create a comment. Comments may be attached to different object types (trace, observation, session, prompt).
method: POST
path: /comments
request: CreateCommentRequest
response: CreateCommentResponse
get:
docs: Get all comments
method: GET
path: /comments
request:
name: GetCommentsRequest
query-parameters:
page:
type: optional<integer>
docs: Page number, starts at 1.
limit:
type: optional<integer>
docs: Limit of items per page. If you encounter api issues due to too large page sizes, try to reduce the limit
objectType:
type: optional<string>
docs: Filter comments by object type (trace, observation, session, prompt).
objectId:
type: optional<string>
docs: Filter comments by object id. If objectType is not provided, an error will be thrown.
authorUserId:
type: optional<string>
docs: Filter comments by author user id.
response: GetCommentsResponse
get-by-id:
docs: Get a comment by id
method: GET
path: /comments/{commentId}
path-parameters:
commentId:
type: string
docs: The unique langfuse identifier of a comment
response: commons.Comment
types:
CreateCommentRequest:
properties:
projectId:
type: string
docs: The id of the project to attach the comment to.
objectType:
type: string
docs: The type of the object to attach the comment to (trace, observation, session, prompt).
objectId:
type: string
docs: The id of the object to attach the comment to. If this does not reference a valid existing object, an error will be thrown.
content:
type: string
docs: The content of the comment. May include markdown. Currently limited to 500 characters.
authorUserId:
type: optional<string>
docs: The id of the user who created the comment.
CreateCommentResponse:
properties:
id:
type: string
docs: The id of the created object in Langfuse
GetCommentsResponse:
properties:
data: list<commons.Comment>
meta: pagination.MetaResponse
+21
View File
@@ -240,6 +240,9 @@ types:
configId:
type: optional<string>
docs: Reference a score config on a score. When set, config and score name must be equal and value must comply to optionally defined numerical range
queueId:
type: optional<string>
docs: Reference an annotation queue on a score. Populated if the score was initially created in an annotation queue.
NumericScore:
extends: BaseScore
properties:
@@ -283,6 +286,18 @@ types:
- double
- string
docs: The value of the score. Must be passed as string for categorical scores, and numeric for boolean and numeric scores
Comment:
properties:
id: string
projectId: string
createdAt: datetime
updatedAt: datetime
objectType: CommentObjectType
objectId: string
content: string
authorUserId: optional<string>
Dataset:
properties:
id: string
@@ -402,6 +417,12 @@ types:
- optional<integer>
- optional<boolean>
- optional<list<string>>
CommentObjectType:
enum:
- TRACE
- OBSERVATION
- SESSION
- PROMPT
DatasetStatus:
enum:
- ACTIVE
+1 -1
View File
@@ -55,7 +55,7 @@ types:
type: string
startDate:
docs: Apply only to generations which are newer than this ISO date.
type: optional<date>
type: optional<datetime>
unit:
docs: Unit used by this model.
type: commons.ModelUsageUnit
+22 -3
View File
@@ -52,10 +52,17 @@ service:
configId:
type: optional<string>
docs: Retrieve only scores with a specific configId.
queueId:
type: optional<string>
docs: Retrieve only scores with a specific annotation queueId.
dataType:
type: optional<commons.ScoreDataType>
docs: Retrieve only scores with a specific dataType.
response: Scores
traceTags:
type: optional<list<string>>
allow-multiple: true
docs: Only scores linked to traces that include all of these tags will be returned.
response: GetScoresResponse
get-by-id:
docs: Get a score
method: GET
@@ -132,7 +139,19 @@ types:
id:
type: string
docs: The id of the created object in Langfuse
Scores:
GetScoresResponseTraceData:
properties:
data: list<commons.Score>
userId:
type: optional<string>
docs: The user ID associated with the trace referenced by score
tags:
type: optional<list<string>>
docs: A list of tags associated with the trace referenced by score
GetScoresResponseData:
properties:
<<: commons.Score
trace: GetScoresResponseTraceData
GetScoresResponse:
properties:
data: list<GetScoresResponseData>
meta: pagination.MetaResponse
+10 -4
View File
@@ -1,6 +1,6 @@
{
"name": "langfuse",
"version": "2.85.0",
"version": "2.89.0",
"author": "engineering@langfuse.com",
"license": "MIT",
"private": true,
@@ -9,15 +9,16 @@
},
"scripts": {
"preinstall": "npx only-allow pnpm",
"infra:dev:up": "docker compose -f ./docker-compose.dev.yml up -d",
"infra:dev:up": "docker compose -f ./docker-compose.dev.yml up -d --wait",
"infra:dev:down": "docker compose -f ./docker-compose.dev.yml down",
"infra:dev:prune": "docker compose -f ./docker-compose.dev.yml down -v",
"db:generate": "turbo run db:generate",
"db:migrate": "turbo run db:migrate",
"db:seed": "turbo run db:seed",
"db:seed:examples": "turbo run db:seed:examples",
"nuke": "bash ./scripts/nuke.sh",
"dx": "pnpm i && pnpm run infra:dev:up && pnpm --filter=shared run db:reset && pnpm --filter=shared run ch:reset && pnpm --filter=shared run db:seed:examples && pnpm run dev",
"dx-f": "pnpm i && pnpm run infra:dev:up && pnpm --filter=shared run db:reset -f && pnpm --filter=shared run ch:reset && pnpm --filter=shared run db:seed:examples && pnpm run dev",
"dx": "pnpm i && pnpm run infra:dev:prune && pnpm run infra:dev:up --pull always && pnpm --filter=shared run db:reset && pnpm --filter=shared run ch:reset && pnpm --filter=shared run db:seed:examples && pnpm run dev",
"dx-f": "pnpm i && pnpm run infra:dev:prune && pnpm run infra:dev:up --pull always && pnpm --filter=shared run db:reset -f && pnpm --filter=shared run ch:reset && pnpm --filter=shared run db:seed:examples && pnpm run dev",
"dx:skip-infra": "pnpm i && pnpm --filter=shared run db:reset && pnpm --filter=shared run ch:reset && pnpm --filter=shared run db:seed:examples && pnpm run dev",
"build": "turbo run build",
"start": "turbo run start",
@@ -80,5 +81,10 @@
}
}
},
"pnpm": {
"overrides": {
"jsonpath-plus": "10.0.0"
}
},
"packageManager": "pnpm@9.5.0"
}
@@ -3,24 +3,30 @@ CREATE TABLE traces (
`timestamp` DateTime64(3),
`name` String,
`user_id` Nullable(String),
`metadata` Map(String, String) CODEC(ZSTD(1)),
`metadata` Map(LowCardinality(String), String),
`release` Nullable(String),
`version` Nullable(String),
`project_id` String,
`public` Bool,
`bookmarked` Bool,
`tags` Array(String),
`input` Nullable(String) CODEC(ZSTD(1)),
`output` Nullable(String) CODEC(ZSTD(1)),
`input` Nullable(String) CODEC(ZSTD(3)),
`output` Nullable(String) CODEC(ZSTD(3)),
`session_id` Nullable(String),
`created_at` DateTime64(3) DEFAULT now(),
`updated_at` DateTime64(3) DEFAULT now(),
updated_at DateTime64(3) DEFAULT now(),
`event_ts` DateTime64(3),
`is_deleted` UInt8,
INDEX idx_id id TYPE bloom_filter(0.001) GRANULARITY 1,
INDEX idx_res_metadata_key mapKeys(metadata) TYPE bloom_filter(0.01) GRANULARITY 1,
INDEX idx_res_metadata_value mapValues(metadata) TYPE bloom_filter(0.01) GRANULARITY 1
) ENGINE = ReplacingMergeTree Partition by toYYYYMM(timestamp)
) ENGINE = ReplacingMergeTree(event_ts, is_deleted) Partition by toYYYYMM(timestamp)
PRIMARY KEY (
project_id,
toDate(timestamp)
)
ORDER BY (
project_id,
toUnixTimestamp(timestamp),
toDate(timestamp),
id
);
);
@@ -7,7 +7,7 @@ CREATE TABLE observations (
`start_time` DateTime64(3),
`end_time` Nullable(DateTime64(3)),
`name` String,
`metadata` Map(LowCardinality(String), String) CODEC(ZSTD(1)),
`metadata` Map(LowCardinality(String), String),
`level` LowCardinality(String),
`status_message` Nullable(String),
`version` Nullable(String),
@@ -16,18 +16,10 @@ CREATE TABLE observations (
`provided_model_name` Nullable(String),
`internal_model_id` Nullable(String),
`model_parameters` Nullable(String),
`provided_input_usage_units` Nullable(Decimal64(12)),
`provided_output_usage_units` Nullable(Decimal64(12)),
`provided_total_usage_units` Nullable(Decimal64(12)),
`input_usage_units` Nullable(Decimal64(12)),
`output_usage_units` Nullable(Decimal64(12)),
`total_usage_units` Nullable(Decimal64(12)),
`unit` Nullable(String),
`provided_input_cost` Nullable(Decimal64(12)),
`provided_output_cost` Nullable(Decimal64(12)),
`provided_total_cost` Nullable(Decimal64(12)),
`input_cost` Nullable(Decimal64(12)),
`output_cost` Nullable(Decimal64(12)),
`provided_usage_details` Map(LowCardinality(String), UInt64),
`usage_details` Map(LowCardinality(String), UInt64),
`provided_cost_details` Map(LowCardinality(String), Decimal64(12)),
`cost_details` Map(LowCardinality(String), Decimal64(12)),
`total_cost` Nullable(Decimal64(12)),
`completion_start_time` Nullable(DateTime64(3)),
`prompt_id` Nullable(String),
@@ -35,16 +27,21 @@ CREATE TABLE observations (
`prompt_version` Nullable(UInt16),
`created_at` DateTime64(3) DEFAULT now(),
`updated_at` DateTime64(3) DEFAULT now(),
event_ts DateTime64(3),
is_deleted UInt8,
INDEX idx_id id TYPE bloom_filter() GRANULARITY 1,
INDEX idx_trace_id trace_id TYPE bloom_filter() GRANULARITY 1,
INDEX idx_project_id project_id TYPE bloom_filter() GRANULARITY 1,
INDEX idx_res_metadata_key mapKeys(metadata) TYPE bloom_filter() GRANULARITY 1,
INDEX idx_res_metadata_value mapValues(metadata) TYPE bloom_filter() GRANULARITY 1
) ENGINE = ReplacingMergeTree Partition by toYYYYMM(start_time)
INDEX idx_project_id project_id TYPE bloom_filter() GRANULARITY 1
) ENGINE = ReplacingMergeTree(event_ts, is_deleted) Partition by toYYYYMM(start_time)
PRIMARY KEY (
project_id,
`type`,
toDate(start_time)
)
ORDER BY (
project_id,
`type`,
trace_id,
toUnixTimestamp(start_time),
toDate(start_time),
id
);
);
@@ -12,14 +12,22 @@ CREATE TABLE scores (
`config_id` Nullable(String),
`data_type` String,
`string_value` Nullable(String),
`queue_id` Nullable(String),
`created_at` DateTime64(3) DEFAULT now(),
`updated_at` DateTime64(3) DEFAULT now(),
event_ts DateTime64(3),
`is_deleted` UInt8,
INDEX idx_id id TYPE bloom_filter(0.001) GRANULARITY 1,
INDEX idx_project_id trace_id TYPE bloom_filter(0.001) GRANULARITY 1
) ENGINE = ReplacingMergeTree Partition by toYYYYMM(timestamp)
INDEX idx_project_trace_observation (project_id, trace_id, observation_id) TYPE bloom_filter(0.001) GRANULARITY 1
) ENGINE = ReplacingMergeTree(event_ts, is_deleted) Partition by toYYYYMM(timestamp)
PRIMARY KEY (
project_id,
toDate(timestamp),
name
)
ORDER BY (
project_id,
trace_id,
toUnixTimestamp(timestamp),
toDate(timestamp),
name,
id
);
)
@@ -0,0 +1 @@
ALTER TABLE observations ADD INDEX IF NOT EXISTS idx_project_id project_id TYPE bloom_filter() GRANULARITY 1;
@@ -0,0 +1 @@
ALTER TABLE observations DROP INDEX IF EXISTS idx_project_id;
@@ -0,0 +1 @@
ALTER TABLE traces DROP INDEX IF EXISTS idx_session_id;
@@ -0,0 +1,2 @@
ALTER TABLE traces ADD INDEX IF NOT EXISTS idx_session_id session_id TYPE bloom_filter() GRANULARITY 1;
ALTER TABLE traces MATERIALIZE INDEX IF EXISTS idx_session_id;
@@ -0,0 +1,25 @@
import { clickhouseClient } from "@langfuse/shared/src/server";
import { prisma } from "../../src/db";
import { redis } from "@langfuse/shared/src/server";
import { prepareClickhouse } from "../../scripts/prepareClickhouse";
async function main() {
try {
const projectIds = ["7a88fb47-b4e2-43b8-a06c-a5ce950dc53a"]; // Example project IDs
await prepareClickhouse(projectIds, {
numberOfDays: 3,
totalObservations: 1000,
});
console.log("Clickhouse preparation completed successfully.");
} catch (error) {
console.error("Error during Clickhouse preparation:", error);
} finally {
await clickhouseClient.close();
await prisma.$disconnect();
redis?.disconnect();
console.log("Disconnected from Clickhouse.");
}
}
main();
+15 -6
View File
@@ -47,16 +47,19 @@
"ch:up": "bash clickhouse/scripts/up.sh",
"ch:down": "bash clickhouse/scripts/down.sh",
"ch:drop": "bash clickhouse/scripts/drop.sh",
"ch:reset": "pnpm run ch:down && pnpm run ch:up"
"ch:reset": "pnpm run ch:down && pnpm run ch:up && pnpm run ch:seed",
"ch:seed": "dotenv -e ../../.env -- tsx clickhouse/scripts/seed.ts",
"load:setup": "dotenv -e ../../.env -- tsx scripts/load-seed.ts"
},
"prisma": {
"seed": "ts-node -r tsconfig-paths/register -r dotenv/config --compiler-options {\"module\":\"CommonJS\"} prisma/seed.ts"
},
"dependencies": {
"@anthropic-ai/tokenizer": "^0.0.4",
"@aws-sdk/client-s3": "^3.627.0",
"@aws-sdk/lib-storage": "^3.667.0",
"@aws-sdk/s3-request-presigner": "^3.554.0",
"@aws-sdk/client-cloudwatch": "^3.675.0",
"@aws-sdk/client-s3": "^3.675.0",
"@aws-sdk/lib-storage": "^3.675.0",
"@aws-sdk/s3-request-presigner": "^3.679.0",
"@clickhouse/client": "^1.4.0",
"@langchain/anthropic": "^0.3.1",
"@langchain/aws": "^0.1.0",
@@ -81,9 +84,9 @@
"nodemailer": "^6.9.15",
"prisma-extension-kysely": "^2.1.0",
"uuid": "^9.0.1",
"winston": "^3.14.2",
"winston": "^3.15.0",
"zod": "^3.23.8",
"zod-to-json-schema": "^3.23.2"
"zod-to-json-schema": "^3.23.5"
},
"devDependencies": {
"@repo/eslint-config": "workspace:*",
@@ -106,11 +109,17 @@
"prisma-kysely": "^1.8.0",
"ts-node": "^10.9.2",
"tsc-watch": "^6.2.0",
"tsx": "^4.19.1",
"typescript": "^5.4.5",
"vitest": "^2.1.2"
},
"peerDependencies": {
"@types/react": "~18.2.79",
"react": "~18.2.0"
},
"pnpm": {
"overrides": {
"jsonpath-plus": "10.0.0"
}
}
}
-4
View File
@@ -1,4 +0,0 @@
{
"trailingComma": "es5",
"printWidth": 120
}
+23 -1
View File
@@ -143,6 +143,18 @@ export type AuditLog = {
before: string | null;
after: string | null;
};
export type BackgroundMigration = {
id: string;
name: string;
script: string;
args: unknown;
state: Generated<unknown>;
finished_at: Timestamp | null;
failed_at: Timestamp | null;
failed_reason: string | null;
worker_id: string | null;
locked_at: Timestamp | null;
};
export type BatchExport = {
id: string;
created_at: Generated<Timestamp>;
@@ -304,7 +316,7 @@ export type Model = {
input_price: string | null;
output_price: string | null;
total_price: string | null;
unit: string;
unit: string | null;
tokenizer_id: string | null;
tokenizer_config: unknown | null;
};
@@ -402,6 +414,14 @@ export type PosthogIntegration = {
enabled: boolean;
created_at: Generated<Timestamp>;
};
export type Price = {
id: string;
created_at: Generated<Timestamp>;
updated_at: Generated<Timestamp>;
model_id: string;
usage_type: string;
price: string;
};
export type Project = {
id: string;
org_id: string;
@@ -546,6 +566,7 @@ export type DB = {
annotation_queues: AnnotationQueue;
api_keys: ApiKey;
audit_logs: AuditLog;
background_migrations: BackgroundMigration;
batch_exports: BatchExport;
comments: Comment;
cron_jobs: CronJobs;
@@ -565,6 +586,7 @@ export type DB = {
organization_memberships: OrganizationMembership;
organizations: Organization;
posthog_integrations: PosthogIntegration;
prices: Price;
project_memberships: ProjectMembership;
projects: Project;
prompts: Prompt;
@@ -0,0 +1,107 @@
-- Temporarily drop observations_view as prompts.config is a calculated column in the view and we need to update it to a JSON type
DROP VIEW IF EXISTS "observations_view"; -- Drop view as column was added in 20240705154048_observation_view_add_created_at_updated_at and update view must have same columns
-- Update prompts.config to JSON type. Use JSON as JSONB is reordering keys for optimization, which is undesired
ALTER TABLE prompts ALTER COLUMN config TYPE JSON USING config::JSON;
ALTER TABLE prompts ALTER COLUMN config SET DEFAULT '{}'::JSON;
-- Recreate the view with the updated prompts.config column
CREATE VIEW "observations_view" AS -- Specify the columns that should be returned in the view, as calculated columns are added but exist in the observations table already
SELECT
o.id,
o.name,
o.start_time,
o.end_time,
o.parent_observation_id,
o.type,
o.trace_id,
o.metadata,
o.model,
o."modelParameters",
o.input,
o.output,
o.level,
o.status_message,
o.completion_start_time,
o.completion_tokens,
o.prompt_tokens,
o.total_tokens,
o.version,
o.project_id,
o.created_at,
o.updated_at,
o.unit,
o.prompt_id,
p.name as prompt_name, -- added in this change
p.version as prompt_version, -- added in this change
o.input_cost,
o.output_cost,
o.total_cost,
o.internal_model,
m.id AS "model_id",
m.start_date AS "model_start_date",
m.input_price,
m.output_price,
m.total_price,
m.tokenizer_config AS "tokenizer_config",
CASE
WHEN o.calculated_input_cost IS NULL AND o.input_cost IS NULL AND o.output_cost IS NULL AND o.total_cost IS NULL THEN
o.prompt_tokens::decimal * m.input_price
ELSE
COALESCE(o.calculated_input_cost, o.input_cost)
END AS "calculated_input_cost",
CASE
WHEN o.calculated_output_cost IS NULL AND o.input_cost IS NULL AND o.output_cost IS NULL AND o.total_cost IS NULL THEN
o.completion_tokens::decimal * m.output_price
ELSE
COALESCE(o.calculated_output_cost, o.output_cost)
END AS "calculated_output_cost",
CASE
WHEN o.calculated_total_cost IS NULL AND o.input_cost IS NULL AND o.output_cost IS NULL AND o.total_cost IS NULL THEN
CASE
WHEN m.total_price IS NOT NULL AND o.total_tokens IS NOT NULL THEN
m.total_price * o.total_tokens
ELSE
o.prompt_tokens::decimal * m.input_price +
o.completion_tokens::decimal * m.output_price
END
ELSE
COALESCE(o.calculated_total_cost, o.total_cost)
END AS "calculated_total_cost",
CASE WHEN o.end_time IS NULL THEN NULL ELSE (EXTRACT(EPOCH FROM o."end_time") - EXTRACT(EPOCH FROM o."start_time"))::double precision END AS "latency",
CASE WHEN o.completion_start_time IS NOT NULL AND o.start_time IS NOT NULL THEN EXTRACT(EPOCH FROM (completion_start_time - start_time))::double precision ELSE NULL END as "time_to_first_token"
FROM
observations o
LEFT JOIN LATERAL (
SELECT
models.*
FROM
models
WHERE (models.project_id = o.project_id OR models.project_id IS NULL)
AND models.model_name = o.internal_model
AND (models.start_date < o.start_time OR models.start_date IS NULL)
AND o.unit::TEXT = models.unit
ORDER BY
models.project_id ASC, -- in postgres, NULLs are sorted last when ordering ASC
models.start_date DESC NULLS LAST -- now, NULLs are sorted last when ordering DESC as well
LIMIT 1
) m ON TRUE
LEFT JOIN LATERAL (
SELECT
prompts.*
FROM
prompts
WHERE prompts.id = o.prompt_id
AND prompts.project_id = o.project_id
LIMIT 1
) p ON TRUE
-- requirements:
-- 1. The view should return all columns from the observations table
-- 2. The view should match with only one model for each observation if:
-- a. The model has the same project_id as the observation, otherwise the model without project_id.
-- b. The model has the same model_name as the observation
-- c. The model has a start_date that is less than the observation start_time, otherwise the model without start_date
-- d. The model has the same unit as the observation
@@ -0,0 +1,18 @@
INSERT INTO models (
id,
project_id,
model_name,
match_pattern,
start_date,
input_price,
output_price,
total_price,
unit,
tokenizer_id,
tokenizer_config
)
VALUES
-- https://docs.anthropic.com/en/docs/about-claude/models#model-comparison-table
('cm2krz1uf000208jjg5653iud', NULL, 'claude-3.5-haiku-20241022', '(?i)^(claude-3-5-sonnet-20241022|anthropic\.claude-3-5-sonnet-20241022-v2:0|claude-3-5-sonnet-V2@20241022)$', NULL, 0.000003, 0.000015, NULL, 'TOKENS', 'claude', NULL),
('cm2ks2vzn000308jjh4ze1w7q', NULL, 'claude-3.5-haiku-latest', '(?i)^(claude-3-5-sonnet-latest)$', NULL, 0.000003, 0.000015, NULL, 'TOKENS', 'claude', NULL)
@@ -0,0 +1,3 @@
-- https://docs.anthropic.com/en/docs/about-claude/models#model-comparison-table
UPDATE models SET model_name = 'claude-3.5-sonnet-20241022' WHERE id = 'cm2krz1uf000208jjg5653iud';
UPDATE models SET model_name = 'claude-3.5-sonnet-latest' WHERE id = 'cm2ks2vzn000308jjh4ze1w7q';
@@ -0,0 +1,62 @@
-- CreateTable
CREATE TABLE "prices" (
"id" TEXT NOT NULL,
"created_at" TIMESTAMP(3) NOT NULL DEFAULT CURRENT_TIMESTAMP,
"updated_at" TIMESTAMP(3) NOT NULL DEFAULT CURRENT_TIMESTAMP,
"model_id" TEXT NOT NULL,
"usage_type" TEXT NOT NULL,
"price" DECIMAL(65,30) NOT NULL,
CONSTRAINT "prices_pkey" PRIMARY KEY ("id")
);
-- CreateIndex
CREATE INDEX "prices_model_id_idx" ON "prices"("model_id");
-- CreateIndex
CREATE UNIQUE INDEX "prices_model_id_usage_type_key" ON "prices"("model_id", "usage_type");
-- AddForeignKey
ALTER TABLE "prices" ADD CONSTRAINT "prices_model_id_fkey" FOREIGN KEY ("model_id") REFERENCES "models"("id") ON DELETE CASCADE ON UPDATE CASCADE;
-- Alter Table to make unit column nullable
ALTER TABLE "models" ALTER COLUMN "unit" DROP NOT NULL;
-- Create a temporary table to store the new prices for user-defined models
CREATE TEMPORARY TABLE temp_prices AS
SELECT
id AS model_id,
'input' AS usage_type,
input_price AS price
FROM models
WHERE project_id IS NOT NULL AND input_price IS NOT NULL
UNION ALL
SELECT
id AS model_id,
'output' AS usage_type,
output_price AS price
FROM models
WHERE project_id IS NOT NULL AND output_price IS NOT NULL
UNION ALL
SELECT
id AS model_id,
'total' AS usage_type,
total_price AS price
FROM models
WHERE project_id IS NOT NULL AND total_price IS NOT NULL;
-- Insert into prices table
INSERT INTO prices (id, created_at, updated_at, model_id, usage_type, price)
SELECT
md5(random()::text || clock_timestamp()::text || model_id::text || usage_type::text)::uuid AS id,
NOW() AS created_at,
NOW() AS updated_at,
model_id,
usage_type,
price
FROM temp_prices
ON CONFLICT (model_id, usage_type) DO NOTHING;
-- Drop the temporary table
DROP TABLE temp_prices;
@@ -0,0 +1,18 @@
-- CreateTable
CREATE TABLE "background_migrations" (
"id" TEXT NOT NULL,
"name" TEXT NOT NULL,
"script" TEXT NOT NULL,
"args" JSONB NOT NULL,
"finished_at" TIMESTAMP(3),
"failed_at" TIMESTAMP(3),
"failed_reason" TEXT,
"worker_id" TEXT,
"locked_at" TIMESTAMP(3),
CONSTRAINT "background_migrations_pkey" PRIMARY KEY ("id")
);
-- CreateIndex
CREATE UNIQUE INDEX "background_migrations_name_key" ON "background_migrations"("name");
@@ -0,0 +1,2 @@
INSERT INTO background_migrations (id, name, script, args)
VALUES ('32859a35-98f5-4a4a-b438-ebc579349e00', '20241024_1216_add_generations_cost_backfill', 'addGenerationsCostBackfill', '{}');
@@ -0,0 +1,2 @@
INSERT INTO background_migrations (id, name, script, args)
VALUES ('5960f22a-748f-480c-b2f3-bc4f9d5d84bc', '20241024_1730_migrate_traces_from_pg_to_ch', 'migrateTracesFromPostgresToClickhouse', '{}');
@@ -0,0 +1,2 @@
INSERT INTO background_migrations (id, name, script, args)
VALUES ('7526e7c9-0026-4595-af2c-369dfd9176ec', '20241024_1737_migrate_observations_from_pg_to_ch', 'migrateObservationsFromPostgresToClickhouse', '{}');
@@ -0,0 +1,2 @@
INSERT INTO background_migrations (id, name, script, args)
VALUES ('94e50334-50d3-4e49-ad2e-9f6d92c85ef7', '20241024_1738_migrate_scores_from_pg_to_ch', 'migrateScoresFromPostgresToClickhouse', '{}');
@@ -0,0 +1,2 @@
-- DropIndex
DROP INDEX CONCURRENTLY "prices_model_id_idx";
@@ -0,0 +1,2 @@
-- AlterTable
ALTER TABLE "background_migrations" ADD COLUMN "state" jsonb NOT NULL DEFAULT '{}';
@@ -0,0 +1,29 @@
INSERT INTO models (
id,
project_id,
model_name,
match_pattern,
start_date,
input_price,
output_price,
total_price,
unit,
tokenizer_id,
tokenizer_config
)
VALUES
-- https://docs.anthropic.com/en/docs/about-claude/models#model-comparison-table
('cm34aq60d000207ml0j1h31ar', NULL, 'claude-3-5-haiku-20241022', '(?i)^(claude-3-5-haiku-20241022|anthropic\.claude-3-5-haiku-20241022-v1:0|claude-3-5-haiku-V1@20241022)$', NULL, 0.000001, 0.000005, NULL, 'TOKENS', 'claude', NULL),
('cm34aqb9h000307ml6nypd618', NULL, 'claude-3.5-haiku-latest', '(?i)^(claude-3-5-haiku-latest)$', NULL, 0.000001, 0.000005, NULL, 'TOKENS', 'claude', NULL);
INSERT INTO prices (
id,
model_id,
usage_type,
price
)
VALUES
('cm34ax6mc000008jkfqed92mb', 'cm34aq60d000207ml0j1h31ar', 'input', 0.000001),
('cm34axb2o000108jk09wn9b47', 'cm34aqb9h000307ml6nypd618', 'input', 0.000001),
('cm34axeie000208jk8b2ke2t8', 'cm34aq60d000207ml0j1h31ar', 'output', 0.000005),
('cm34axi67000308jk7x1a7qko', 'cm34aqb9h000307ml6nypd618', 'output', 0.000005);
File diff suppressed because it is too large Load Diff
+41 -41
View File
@@ -256,7 +256,7 @@ async function main() {
const queueIds = await generateQueuesForProject(
[project1, project2],
configIdsAndNames
configIdsAndNames,
);
const promptIds = await generatePromptsForProject([project1, project2]);
@@ -282,11 +282,11 @@ async function main() {
project2,
promptIds,
queueIds,
configIdsAndNames
configIdsAndNames,
);
logger.info(
`Seeding ${traces.length} traces, ${observations.length} observations, and ${scores.length} scores`
`Seeding ${traces.length} traces, ${observations.length} observations, and ${scores.length} scores`,
);
await uploadObjects(
@@ -296,7 +296,7 @@ async function main() {
sessions,
events,
comments,
queueItems
queueItems,
);
// If openai key is in environment, add it to the projects LLM API keys
@@ -314,7 +314,7 @@ async function main() {
});
} else {
logger.warn(
"No OPENAI_API_KEY found in environment. Skipping seeding LLM API key."
"No OPENAI_API_KEY found in environment. Skipping seeding LLM API key.",
);
}
@@ -449,7 +449,7 @@ async function main() {
for (const datasetItemId of datasetItemIds) {
const relevantObservations = observations.filter(
(o) => o.projectId === project2.id
(o) => o.projectId === project2.id,
);
const observation =
relevantObservations[
@@ -492,7 +492,7 @@ async function uploadObjects(
sessions: Prisma.TraceSessionCreateManyInput[],
events: Prisma.ObservationCreateManyInput[],
comments: Prisma.CommentCreateManyInput[],
queueItems: Prisma.AnnotationQueueItemCreateManyInput[]
queueItems: Prisma.AnnotationQueueItemCreateManyInput[],
) {
let promises: Prisma.PrismaPromise<unknown>[] = [];
@@ -506,14 +506,14 @@ async function uploadObjects(
},
create: chunk[0]!,
update: {},
})
}),
);
});
for (let i = 0; i < promises.length; i++) {
if (i + 1 >= promises.length || i % Math.ceil(promises.length / 10) === 0)
logger.info(
`Seeding of Sessions ${((i + 1) / promises.length) * 100}% complete`
`Seeding of Sessions ${((i + 1) / promises.length) * 100}% complete`,
);
await promises[i];
}
@@ -524,13 +524,13 @@ async function uploadObjects(
promises.push(
prisma.trace.createMany({
data: chunk,
})
}),
);
});
for (let i = 0; i < promises.length; i++) {
if (i + 1 >= promises.length || i % Math.ceil(promises.length / 10) === 0)
logger.info(
`Seeding of Traces ${((i + 1) / promises.length) * 100}% complete`
`Seeding of Traces ${((i + 1) / promises.length) * 100}% complete`,
);
await promises[i];
}
@@ -540,14 +540,14 @@ async function uploadObjects(
promises.push(
prisma.observation.createMany({
data: chunk,
})
}),
);
});
for (let i = 0; i < promises.length; i++) {
if (i + 1 >= promises.length || i % Math.ceil(promises.length / 10) === 0)
logger.info(
`Seeding of Observations ${((i + 1) / promises.length) * 100}% complete`
`Seeding of Observations ${((i + 1) / promises.length) * 100}% complete`,
);
await promises[i];
}
@@ -557,14 +557,14 @@ async function uploadObjects(
promises.push(
prisma.observation.createMany({
data: chunk,
})
}),
);
});
for (let i = 0; i < promises.length; i++) {
if (i + 1 >= promises.length || i % Math.ceil(promises.length / 10) === 0)
logger.info(
`Seeding of Events ${((i + 1) / promises.length) * 100}% complete`
`Seeding of Events ${((i + 1) / promises.length) * 100}% complete`,
);
await promises[i];
}
@@ -574,13 +574,13 @@ async function uploadObjects(
promises.push(
prisma.score.createMany({
data: chunk,
})
}),
);
});
for (let i = 0; i < promises.length; i++) {
if (i + 1 >= promises.length || i % Math.ceil(promises.length / 10) === 0)
logger.info(
`Seeding of Scores ${((i + 1) / promises.length) * 100}% complete`
`Seeding of Scores ${((i + 1) / promises.length) * 100}% complete`,
);
await promises[i];
}
@@ -590,13 +590,13 @@ async function uploadObjects(
promises.push(
prisma.comment.createMany({
data: chunk,
})
}),
);
});
for (let i = 0; i < promises.length; i++) {
if (i + 1 >= promises.length || i % Math.ceil(promises.length / 10) === 0)
logger.info(
`Seeding of Comments ${((i + 1) / promises.length) * 100}% complete`
`Seeding of Comments ${((i + 1) / promises.length) * 100}% complete`,
);
await promises[i];
}
@@ -606,13 +606,13 @@ async function uploadObjects(
promises.push(
prisma.annotationQueueItem.createMany({
data: chunk,
})
}),
);
});
for (let i = 0; i < promises.length; i++) {
if (i + 1 >= promises.length || i % Math.ceil(promises.length / 10) === 0)
logger.info(
`Seeding of Annotation Queue Items ${((i + 1) / promises.length) * 100}% complete`
`Seeding of Annotation Queue Items ${((i + 1) / promises.length) * 100}% complete`,
);
await promises[i];
}
@@ -634,7 +634,7 @@ function createObjects(
dataType: ScoreDataType;
categories: ConfigCategory[] | null;
}[]
>
>,
) {
const traces: Prisma.TraceCreateManyInput[] = [];
const observations: Prisma.ObservationCreateManyInput[] = [];
@@ -649,7 +649,7 @@ function createObjects(
// print progress to console with a progress bar that refreshes every 10 iterations
// random date within last 90 days, with a linear bias towards more recent dates
const traceTs = new Date(
Date.now() - Math.floor(Math.random() ** 1.5 * 90 * 24 * 60 * 60 * 1000)
Date.now() - Math.floor(Math.random() ** 1.5 * 90 * 24 * 60 * 60 * 1000),
);
const envTag = envTags[Math.floor(Math.random() * envTags.length)];
@@ -805,11 +805,11 @@ function createObjects(
for (let j = 0; j < Math.floor(Math.random() * 10) + 1; j++) {
// add between 1 and 30 ms to trace timestamp
const spanTsStart = new Date(
traceTs.getTime() + Math.floor(Math.random() * 30)
traceTs.getTime() + Math.floor(Math.random() * 30),
);
// random duration of upto 5000ms
const spanTsEnd = new Date(
spanTsStart.getTime() + Math.floor(Math.random() * 5000)
spanTsStart.getTime() + Math.floor(Math.random() * 5000),
);
const span = {
@@ -844,22 +844,22 @@ function createObjects(
const generationTsStart = new Date(
spanTsStart.getTime() +
Math.floor(
Math.random() * (spanTsEnd.getTime() - spanTsStart.getTime())
)
Math.random() * (spanTsEnd.getTime() - spanTsStart.getTime()),
),
);
const generationTsEnd = new Date(
generationTsStart.getTime() +
Math.floor(
Math.random() *
(spanTsEnd.getTime() - generationTsStart.getTime())
)
(spanTsEnd.getTime() - generationTsStart.getTime()),
),
);
// somewhere in the middle
const generationTsCompletionStart = new Date(
generationTsStart.getTime() +
Math.floor(
(generationTsEnd.getTime() - generationTsStart.getTime()) / 3
)
(generationTsEnd.getTime() - generationTsStart.getTime()) / 3,
),
);
const promptTokens = Math.floor(Math.random() * 1000) + 300;
@@ -880,7 +880,7 @@ function createObjects(
const promptId =
promptIds.get(projectId)![
Math.floor(
Math.random() * Math.floor(promptIds.get(projectId)!.length / 2)
Math.random() * Math.floor(promptIds.get(projectId)!.length / 2),
)
];
@@ -962,8 +962,8 @@ function createObjects(
const eventTs = new Date(
spanTsStart.getTime() +
Math.floor(
Math.random() * (spanTsEnd.getTime() - spanTsStart.getTime())
)
Math.random() * (spanTsEnd.getTime() - spanTsStart.getTime()),
),
);
events.push({
@@ -985,7 +985,7 @@ function createObjects(
}
// find unique sessions by id and projectid
const uniqueSessions: Prisma.TraceSessionCreateManyInput[] = Array.from(
new Set(sessions.map((session) => JSON.stringify(session)))
new Set(sessions.map((session) => JSON.stringify(session))),
).map((session) => JSON.parse(session) as Prisma.TraceSessionCreateManyInput);
return {
@@ -1007,7 +1007,7 @@ async function generatePromptsForProject(projects: Project[]) {
projects.map(async (project) => {
const promptIdsForProject = await generatePrompts(project);
promptIds.set(project.id, promptIdsForProject);
})
}),
);
return promptIds;
}
@@ -1193,7 +1193,7 @@ async function generateConfigsForProject(projects: Project[]) {
projects.map(async (project) => {
const configNameAndId = await generateConfigs(project);
projectIdsToConfigs.set(project.id, configNameAndId);
})
}),
);
return projectIdsToConfigs;
}
@@ -1343,7 +1343,7 @@ async function generateQueuesForProject(
dataType: ScoreDataType;
categories: ConfigCategory[] | null;
}[]
>
>,
) {
const projectIdsToQueues: Map<string, string[]> = new Map();
@@ -1351,10 +1351,10 @@ async function generateQueuesForProject(
projects.map(async (project) => {
const queueIds = await generateQueues(
project,
configIdsAndNames.get(project.id) ?? []
configIdsAndNames.get(project.id) ?? [],
);
projectIdsToQueues.set(project.id, queueIds);
})
}),
);
return projectIdsToQueues;
}
@@ -1366,7 +1366,7 @@ async function generateQueues(
id: string;
dataType: ScoreDataType;
categories: ConfigCategory[] | null;
}[]
}[],
) {
const queue = {
id: `queue-${v4()}`,
+138
View File
@@ -0,0 +1,138 @@
import { randomUUID } from "crypto";
import { prisma } from "../src/db";
import {
clickhouseClient,
getDisplaySecretKey,
hashSecretKey,
logger,
} from "../src/server";
import { prepareClickhouse } from "./prepareClickhouse";
import { redis } from "../src/server";
const createRandomProjectId = () => randomUUID().toString();
const prepareProjectsAndApiKeys = async (
numOfProjects: number,
opts: { requiredProjectIds: string[] },
) => {
const { requiredProjectIds } = opts;
const projectsToCreate = numOfProjects - requiredProjectIds.length;
const projectIds = [...requiredProjectIds];
for (let i = 0; i < projectsToCreate; i++) {
projectIds.push(createRandomProjectId());
}
const operations = projectIds.map(async (projectId) => {
const orgId = `org-${projectId}`;
await prisma.organization.upsert({
where: { id: orgId },
update: {},
create: {
id: orgId,
name: `Organization for ${projectId}`,
},
});
await prisma.project.upsert({
where: { id: projectId },
update: {},
create: {
id: projectId,
name: `Project ${projectId}`,
orgId: orgId,
},
});
const apiKeyId = `api-key-${projectId}`;
const apiKeyExists = await prisma.apiKey.findUnique({
where: { id: apiKeyId },
});
if (!apiKeyExists) {
const sk = await hashSecretKey(
`sk-${Math.random().toString(36).substr(2, 9)}`,
);
await prisma.apiKey.create({
data: {
id: apiKeyId,
note: `API Key for ${projectId}`,
publicKey: `pk-${Math.random().toString(36).substr(2, 9)}`,
hashedSecretKey: sk,
displaySecretKey: getDisplaySecretKey(sk),
project: {
connect: {
id: projectId,
},
},
},
});
}
});
await Promise.all(operations);
return projectIds;
};
async function main() {
let numOfProjects = parseInt(process.argv[2], 10);
let numberOfDays = parseInt(process.argv[3], 10);
let totalObservations = parseInt(process.argv[4], 10);
logger.info(process.argv);
logger.info(
`Preparing Clickhouse for ${numOfProjects} projects and ${numberOfDays} days with max Observations ${totalObservations}.`,
);
if (isNaN(totalObservations)) {
logger.warn(
"Total observations not provided or invalid. Defaulting to 1000 observations.",
);
totalObservations = 1000;
}
if (isNaN(numOfProjects)) {
logger.warn(
"Number of projects not provided or invalid. Defaulting to 10 projects.",
);
numOfProjects = 10;
}
if (isNaN(numberOfDays)) {
logger.warn(
"Number of days not provided or invalid. Defaulting to 3 days.",
);
numberOfDays = 3;
}
logger.info(
`Preparing Clickhouse for ${numOfProjects} projects and ${numberOfDays} days with max Observations ${totalObservations}.`,
);
try {
const projectIds = [
"7a88fb47-b4e2-43b8-a06c-a5ce950dc53a",
"239ad00f-562f-411d-af14-831c75ddd875",
];
const createdProjectIds = await prepareProjectsAndApiKeys(numOfProjects, {
requiredProjectIds: projectIds,
});
await prepareClickhouse(createdProjectIds, {
numberOfDays,
totalObservations: totalObservations ?? 1000,
});
logger.info("Clickhouse preparation completed successfully.");
} catch (error) {
logger.error("Error during Clickhouse preparation:", error);
} finally {
await clickhouseClient.close();
await prisma.$disconnect();
redis?.disconnect();
logger.info("Disconnected from Clickhouse.");
}
}
main();
@@ -0,0 +1,247 @@
import { prisma } from "../src/db";
import { clickhouseClient, logger } from "../src/server";
function randn_bm(min: number, max: number, skew: number) {
let u = 0,
v = 0;
while (u === 0) u = Math.random(); //Converting [0,1) to (0,1)
while (v === 0) v = Math.random();
let num = Math.sqrt(-2.0 * Math.log(u)) * Math.cos(2.0 * Math.PI * v);
num = num / 10.0 + 0.5; // Translate to 0 -> 1
if (num > 1 || num < 0)
num = randn_bm(min, max, skew); // resample between 0 and 1 if out of range
else {
num = Math.pow(num, skew); // Skew
num *= max - min; // Stretch to fill range
num += min; // offset to min
}
return num;
}
export const prepareClickhouse = async (
projectIds: string[],
opts: {
numberOfDays: number;
totalObservations: number;
},
) => {
logger.info(
`Preparing Clickhouse for ${projectIds.length} projects and ${opts.numberOfDays} days.`,
);
const projectData = projectIds.map((projectId) => {
const observationsPerProject = Math.ceil(
randn_bm(0, opts.totalObservations, 2),
); // Skew the number of observations
const tracesPerProject = Math.floor(observationsPerProject / 6); // On average, one trace should have 6 observations
const scoresPerProject = tracesPerProject * 10; // On average, one trace should have 10 scores
return {
projectId,
observationsPerProject,
tracesPerProject,
scoresPerProject,
};
});
for (const data of projectData) {
const {
projectId,
tracesPerProject,
observationsPerProject,
scoresPerProject,
} = data;
logger.info(
`Preparing Clickhouse for ${projectId}: Traces: ${tracesPerProject}, Scores: ${scoresPerProject}, Observations: ${observationsPerProject}`,
);
const tracesQuery = `
INSERT INTO traces
SELECT toString(number) AS id,
toDateTime(now() - randUniform(0, ${opts.numberOfDays} * 24 * 60 * 60)) AS timestamp,
concat('name_', toString(rand() % 100)) AS name,
concat('user_id_', toInt64(randExponential(1 / 100))) AS user_id,
map('key', 'value') AS metadata,
concat('release_', toString(randUniform(0, 100))) AS release,
concat('version_', toString(randUniform(0, 100))) AS version,
'${projectId}' AS project_id,
if(rand() < 0.8, true, false) as public,
if(rand() < 0.8, true, false) as bookmarked,
array('tag1', 'tag2') as tags,
repeat('input', toInt64(randExponential(1 / 100))) AS input,
repeat('output', toInt64(randExponential(1 / 100))) AS output,
if(randUniform(0, 1) < 0.2, NULL, concat('session_', toString(rand() % 1000))) AS session_id,
timestamp AS created_at,
timestamp AS updated_at,
timestamp AS event_ts,
0 AS is_deleted
FROM numbers(${tracesPerProject});
`;
const observationsQuery = `
INSERT INTO observations
SELECT toString(number) AS id,
toString(floor(randUniform(0, ${tracesPerProject}))) AS trace_id,
'${projectId}' AS project_id,
if(randUniform(0, 1) < 0.47, 'GENERATION', if(randUniform(0, 1) < 0.94, 'SPAN', 'EVENT')) AS type,
toString(rand()) AS parent_observation_id,
toDateTime(now() - randUniform(0, ${opts.numberOfDays} * 24 * 60 * 60)) AS start_time,
addSeconds(start_time, if(rand() < 0.6, floor(randUniform(0, 20)), floor(randUniform(0, 3600)))) AS end_time,
concat('name', toString(rand() % 100)) AS name,
map('key', 'value') AS metadata,
if(randUniform(0, 1) < 0.9, 'DEFAULT', if(randUniform(0, 1) < 0.5, 'ERROR', if(randUniform(0, 1) < 0.5, 'DEBUG', 'WARNING'))) AS level,
'status_message' AS status_message,
'version' AS version,
repeat('input', toInt64(randExponential(1 / 100))) AS input,
repeat('output', toInt64(randExponential(1 / 100))) AS output,
case
when number % 2 = 0 then 'claude-3-haiku-20240307'
else 'gpt-4'
end as provided_model_name,
case
when number % 2 = 0 then 'cltr0w45b000008k1407o9qv1'
else 'clrntkjgy000f08jx79v9g1xj'
end as internal_model_id,
'{"temperature": 0.7, "max_tokens": 150}' AS model_parameters,
map('input', toUInt64(randUniform(0, 1000)), 'output', toUInt64(randUniform(0, 1000)), 'total', toUInt64(randUniform(0, 2000))) AS provided_usage_details,
map('input', toUInt64(randUniform(0, 1000)), 'output', toUInt64(randUniform(0, 1000)), 'total', toUInt64(randUniform(0, 2000))) AS usage_details,
map('input', toDecimal64(randUniform(0, 1000), 12), 'output', toDecimal64(randUniform(0, 1000), 12), 'total', toDecimal64(randUniform(0, 2000), 12)) AS provided_cost_details,
map('input', toDecimal64(randUniform(0, 1000), 12), 'output', toDecimal64(randUniform(0, 1000), 12), 'total', toDecimal64(randUniform(0, 2000), 12)) AS cost_details,
toDecimal64(randUniform(0, 2000), 12) AS total_cost,
start_time AS completion_start_time,
toString(rand()) AS prompt_id,
toString(rand()) AS prompt_name,
1000 AS prompt_version,
start_time AS created_at,
start_time AS updated_at,
start_time AS event_ts,
0 AS is_deleted
FROM numbers(${observationsPerProject});
`;
const scoresQuery = `
INSERT INTO scores
SELECT toString(floor(randUniform(0, 100))) AS id,
toDateTime(now() - randUniform(0, ${opts.numberOfDays} * 24 * 60 * 60)) AS timestamp,
'${projectId}' AS project_id,
toString(floor(randUniform(0, ${tracesPerProject}))) AS trace_id,
if(
rand() > 0.9,
toString(floor(randUniform(0, ${observationsPerProject}))),
NULL
) AS observation_id,
concat('name_', toString(rand() % 10)) AS name,
randUniform(0, 100) as value,
'API' as source,
'comment' as comment,
toString(rand() % 100) as author_user_id,
toString(rand() % 100) as config_id,
if (rand() < 0.33, 'NUMERIC', if (rand() < 0.5, 'CATEGORICAL', 'BOOLEAN')) as data_type,
toString(rand() % 100) as string_value,
NULL as queue_id,
timestamp AS created_at,
timestamp AS updated_at,
timestamp AS event_ts,
0 AS is_deleted
FROM numbers(${scoresPerProject});
`;
const queries = [tracesQuery, scoresQuery, observationsQuery];
for (const query of queries) {
logger.info(`Executing query: ${query}`);
await clickhouseClient.command({
query,
clickhouse_settings: {
wait_end_of_query: 1,
},
});
}
// we also need to upsert trace sessions in postgres
const sessionQuery = `
SELECT session_id, project_id
FROM traces
WHERE session_id IS NOT NULL;
`;
const sessionResult = await clickhouseClient.query({
query: sessionQuery,
format: "JSONEachRow",
});
const sessionData = await sessionResult.json<{
session_id: string;
project_id: string;
}>();
const idProjectIdCombinations = sessionData.map((session) => ({
id: session.session_id,
projectId: session.project_id,
public: Math.random() < 0.1,
bookmarked: Math.random() < 0.1,
}));
await prisma.traceSession.createMany({
data: idProjectIdCombinations,
skipDuplicates: true,
});
}
const tables = ["traces", "scores", "observations"];
for (const table of tables) {
const query = `
SELECT
project_id,
count() AS per_project_count,
bar(per_project_count, 0, (
SELECT count(*)
FROM ${table}
), 50) AS bar_representation
FROM ${table}
GROUP BY project_id
ORDER BY count() desc
`;
const result = await clickhouseClient.query({
query,
format: "TabSeparated",
});
logger.info(
`${table.charAt(0).toUpperCase() + table.slice(1)} per Project: \n` +
(await result.text()),
);
}
const tablesWithDateColumns = [
{ name: "traces", dateColumn: "timestamp" },
{ name: "scores", dateColumn: "timestamp" },
{ name: "observations", dateColumn: "start_time" },
];
for (const { name: table, dateColumn } of tablesWithDateColumns) {
const query = `
SELECT
toDate(${dateColumn}) AS event_date,
count() AS per_date_count,
bar(per_date_count, 0, (
SELECT count(*)
FROM ${table}
), 50) AS bar_representation
FROM ${table}
GROUP BY event_date
ORDER BY event_date desc
`;
const result = await clickhouseClient.query({
query,
format: "TabSeparated",
});
logger.info(
`${table.charAt(0).toUpperCase() + table.slice(1)} per Date: \n` +
(await result.text()),
);
}
};
+33 -30
View File
@@ -21,37 +21,40 @@ export class PrismaClientSingleton {
return PrismaClientSingleton.instance;
}
const client = new PrismaClient<
Prisma.PrismaClientOptions,
"warn" | "error" | "query"
>({
log: [
{ emit: "event", level: "query" },
{ emit: "event", level: "error" },
{ emit: "event", level: "warn" },
],
});
if (env.NODE_ENV === "development") {
client.$on("query", (event) => {
logger.info(`prisma:query ${event.query}, ${event.duration}ms`);
});
}
client.$on("warn", (event) => {
logger.warn(`prisma:warn ${event.message}`);
});
client.$on("error", (event) => {
logger.error(`prisma:error ${event.message}`);
});
PrismaClientSingleton.instance = client;
PrismaClientSingleton.instance = createPrismaInstance();
return PrismaClientSingleton.instance;
}
}
const createPrismaInstance = () => {
const client = new PrismaClient<
Prisma.PrismaClientOptions,
"warn" | "error" | "query"
>({
log: [
{ emit: "event", level: "query" },
{ emit: "event", level: "error" },
{ emit: "event", level: "warn" },
],
});
if (env.NODE_ENV === "development") {
client.$on("query", (event) => {
logger.info(`prisma:query ${event.query}, ${event.duration}ms`);
});
}
client.$on("warn", (event) => {
logger.warn(`prisma:warn ${event.message}`);
});
client.$on("error", (event) => {
logger.error(`prisma:error ${event.message}`);
});
return client;
};
export class KyselySingleton {
private static instance: { $kysely: Kysely<DB> };
@@ -73,7 +76,7 @@ export class KyselySingleton {
createQueryCompiler: () => new PostgresQueryCompiler(),
},
}),
})
}),
);
return KyselySingleton.instance;
@@ -86,7 +89,7 @@ declare const globalThis: {
} & typeof global;
if (process.env.NODE_ENV === "development") {
globalThis.prismaGlobal ??= new PrismaClient(); // regular instantiation
globalThis.prismaGlobal ??= createPrismaInstance(); // regular instantiation
globalThis.kyselyPrismaGlobal ??= globalThis.prismaGlobal.$extends(
kyselyExtension({
kysely: (driver) =>
@@ -100,9 +103,9 @@ if (process.env.NODE_ENV === "development") {
createQueryCompiler: () => new PostgresQueryCompiler(),
},
}),
})
}),
);
};
}
export const prisma =
globalThis.prismaGlobal ?? PrismaClientSingleton.getInstance();
+14 -10
View File
@@ -22,7 +22,7 @@ const EnvSchema = z.object({
.string()
.length(
64,
"ENCRYPTION_KEY must be 256 bits, 64 string characters in hex format, generate via: openssl rand -hex 32"
"ENCRYPTION_KEY must be 256 bits, 64 string characters in hex format, generate via: openssl rand -hex 32",
)
.optional(),
LANGFUSE_CACHE_PROMPT_ENABLED: z.enum(["true", "false"]).default("false"),
@@ -30,20 +30,24 @@ const EnvSchema = z.object({
CLICKHOUSE_URL: z.string().url().optional(),
CLICKHOUSE_USER: z.string().optional(),
CLICKHOUSE_PASSWORD: z.string().optional(),
LANGFUSE_INGESTION_FLUSH_DELAY_MS: z.coerce
.number()
.nonnegative()
.default(10000),
LANGFUSE_INGESTION_FLUSH_ATTEMPTS: z.coerce.number().positive().default(3),
LANGFUSE_INGESTION_BUFFER_TTL_SECONDS: z.coerce
.number()
.positive()
.default(60 * 10),
SALT: z.string().optional(), // used by components imported by web package
LANGFUSE_LOG_LEVEL: z
.enum(["trace", "debug", "info", "warn", "error", "fatal"])
.optional(),
LANGFUSE_LOG_FORMAT: z.enum(["text", "json"]).default("text"),
ENABLE_AWS_CLOUDWATCH_METRIC_PUBLISHING: z
.enum(["true", "false"])
.default("false"),
LANGFUSE_S3_EVENT_UPLOAD_ENABLED: z.enum(["true", "false"]).default("false"),
LANGFUSE_S3_EVENT_UPLOAD_BUCKET: z.string().optional(),
LANGFUSE_S3_EVENT_UPLOAD_PREFIX: z.string().default(""),
LANGFUSE_S3_EVENT_UPLOAD_REGION: z.string().optional(),
LANGFUSE_S3_EVENT_UPLOAD_ENDPOINT: z.string().optional(),
LANGFUSE_S3_EVENT_UPLOAD_ACCESS_KEY_ID: z.string().optional(),
LANGFUSE_S3_EVENT_UPLOAD_SECRET_ACCESS_KEY: z.string().optional(),
LANGFUSE_S3_EVENT_UPLOAD_FORCE_PATH_STYLE: z
.enum(["true", "false"])
.default("false"),
});
export const env = EnvSchema.parse(removeEmptyEnvVariables(process.env));
@@ -0,0 +1,3 @@
export * from "./scoreTypes";
export * from "./types";
export * from "./scoreConfigTypes";
@@ -55,6 +55,7 @@ const ScoreBase = z.object({
configId: z.string().nullish(),
createdAt: z.coerce.date(),
updatedAt: z.coerce.date(),
queueId: z.string().nullish(),
});
const BaseScoreBody = z.object({
@@ -83,13 +84,13 @@ export const ScoreBodyWithoutConfig = z.discriminatedUnion("dataType", [
z.object({
value: z.number(),
dataType: z.literal("NUMERIC"),
}),
})
),
BaseScoreBody.merge(
z.object({
value: z.string(),
dataType: z.literal("CATEGORICAL"),
}),
})
),
BaseScoreBody.merge(
z.object({
@@ -97,7 +98,7 @@ export const ScoreBodyWithoutConfig = z.discriminatedUnion("dataType", [
message: "Value must be either 0 or 1",
}),
dataType: z.literal("BOOLEAN"),
}),
})
),
]);
@@ -161,7 +162,7 @@ export const ScorePropsAgainstConfig = z.union([
*/
export const filterAndValidateDbScoreList = (
scores: Score[],
onParseError?: (error: z.ZodError) => void,
onParseError?: (error: z.ZodError) => void
): APIScore[] =>
scores.reduce((acc, ts) => {
const result = APIScoreSchema.safeParse(ts);
@@ -198,14 +199,14 @@ export const PostScoresBody = z.discriminatedUnion("dataType", [
value: z.number(),
dataType: z.literal("NUMERIC"),
configId: z.string().nullish(),
}),
})
),
BaseScoreBody.merge(
z.object({
value: z.string(),
dataType: z.literal("CATEGORICAL"),
configId: z.string().nullish(),
}),
})
),
BaseScoreBody.merge(
z.object({
@@ -215,14 +216,14 @@ export const PostScoresBody = z.discriminatedUnion("dataType", [
}),
dataType: z.literal("BOOLEAN"),
configId: z.string().nullish(),
}),
})
),
BaseScoreBody.merge(
z.object({
value: z.union([z.string(), z.number()]),
dataType: z.undefined(),
configId: z.string().nullish(),
}),
})
),
]);
@@ -234,6 +235,8 @@ export const GetScoresQuery = z.object({
userId: z.string().nullish(),
dataType: z.enum(ScoreDataType).nullish(),
configId: z.string().nullish(),
queueId: z.string().nullish(),
traceTags: z.union([z.array(z.string()), z.string()]).nullish(),
name: z.string().nullish(),
fromTimestamp: stringDateTime,
toTimestamp: stringDateTime,
@@ -255,8 +258,9 @@ const LegacyGetScoreResponseDataV1 = z.intersection(
z.object({
trace: z.object({
userId: z.string().nullish(),
tags: z.array(z.string()).nullish(),
}),
}),
})
);
export const GetScoresResponse = z.object({
data: z.array(LegacyGetScoreResponseDataV1),
@@ -265,7 +269,7 @@ export const GetScoresResponse = z.object({
export const legacyFilterAndValidateV1GetScoreList = (
scores: unknown[],
onParseError?: (error: z.ZodError) => void,
onParseError?: (error: z.ZodError) => void
): z.infer<typeof LegacyGetScoreResponseDataV1>[] =>
scores.reduce(
(acc: z.infer<typeof LegacyGetScoreResponseDataV1>[], ts) => {
@@ -278,7 +282,7 @@ export const legacyFilterAndValidateV1GetScoreList = (
}
return acc;
},
[] as z.infer<typeof LegacyGetScoreResponseDataV1>[],
[] as z.infer<typeof LegacyGetScoreResponseDataV1>[]
);
// GET /scores/{scoreId}
@@ -0,0 +1,37 @@
import { ScoreDataType, ScoreSource } from "@prisma/client";
export type CategoricalAggregate = {
type: "CATEGORICAL";
values: string[];
valueCounts: { value: string; count: number }[];
comment?: string | null;
};
export type NumericAggregate = {
type: "NUMERIC";
values: number[];
average: number;
comment?: string | null;
};
export type ScoreAggregate = Record<
string,
CategoricalAggregate | NumericAggregate
>;
export type ScoreSimplified = {
name: string;
dataType: ScoreDataType;
source: ScoreSource;
value?: number | null;
comment?: string | null;
stringValue?: string | null;
};
export type LastUserScore = ScoreSimplified & {
timestamp: string;
traceId: string;
observationId?: string | null;
userId: string;
};
+2 -3
View File
@@ -6,7 +6,7 @@ export * from "./interfaces/parseDbOrg";
export * from "./interfaces/customLLMProviderConfigSchemas";
export * from "./tableDefinitions";
export * from "./types";
export * from "./tracesTable";
export * from "./tableDefinitions/tracesTable";
export * from "./server/auth/apiKeys";
export * from "./observationsTable";
export * from "./utils/zod";
@@ -28,8 +28,7 @@ export * from "./features/batchExport/types";
export * from "./features/annotation/types";
// scores
export * from "./features/scores/scoreConfigTypes";
export * from "./features/scores/scoreTypes";
export * from "./features/scores";
// comments
export * from "./features/comments/types";
@@ -1,6 +1,5 @@
import { createClient } from "@clickhouse/client";
import { env } from "../env";
import { env } from "../../env";
export type ClickhouseClientType = ReturnType<typeof createClient>;
@@ -14,3 +13,12 @@ export const clickhouseClient = createClient({
wait_for_async_insert: 1, // if disabled, we won't get errors from clickhouse
},
});
/**
* Accepts a JavaScript date and returns the DateTime in format YYYY-MM-DD HH:MM:SS
*/
export const convertDateToClickhouseDateTime = (date: Date): string => {
// 2024-11-06T20:37:00.000Z -> 2024-11-06 21:37:00
return date.toISOString().slice(0, 19).replace("T", " ");
};
@@ -0,0 +1,7 @@
export const ClickhouseTableNames = {
traces: "traces",
observations: "observations",
scores: "scores",
} as const;
export type ClickhouseTableName = keyof typeof ClickhouseTableNames;
@@ -0,0 +1,37 @@
import { LangfuseNotFoundError } from "../../errors";
import { eventTypes } from "../ingestion/types";
import { ClickhouseTableName, ClickhouseTableNames } from "./schema";
export const isValidTableName = (
tableName: string,
): tableName is ClickhouseTableName =>
Object.keys(ClickhouseTableNames).includes(tableName);
export type IngestionEntityTypes =
| "trace"
| "observation"
| "score"
| "sdk_log";
export const getClickhouseEntityType = (
eventType: string,
): IngestionEntityTypes => {
switch (eventType) {
case eventTypes.TRACE_CREATE:
return "trace";
case eventTypes.OBSERVATION_CREATE:
case eventTypes.OBSERVATION_UPDATE:
case eventTypes.EVENT_CREATE:
case eventTypes.SPAN_CREATE:
case eventTypes.SPAN_UPDATE:
case eventTypes.GENERATION_CREATE:
case eventTypes.GENERATION_UPDATE:
return "observation";
case eventTypes.SCORE_CREATE:
return "score";
case eventTypes.SDK_LOG:
return "sdk_log";
default:
throw new LangfuseNotFoundError(`Unknown event type: ${eventType}`);
}
};
-167
View File
@@ -1,167 +0,0 @@
import z from "zod";
export const clickhouseStringDateSchema = z
.string()
// clickhouse stores UTC like '2024-05-23 18:33:41.602000'
// we need to convert it to '2024-05-23T18:33:41.602000Z'
.transform((str) => str.replace(" ", "T") + "Z")
.pipe(z.string().datetime());
export const observationRecordBaseSchema = z.object({
id: z.string(),
trace_id: z.string().nullish(),
project_id: z.string(),
type: z.string(),
parent_observation_id: z.string().nullish(),
name: z.string().nullish(),
metadata: z.record(z.string()),
level: z.string().nullish(),
status_message: z.string().nullish(),
version: z.string().nullish(),
input: z.string().nullish(),
output: z.string().nullish(),
provided_model_name: z.string().nullish(),
internal_model_id: z.string().nullish(),
model_parameters: z.string().nullish(),
unit: z.string().nullish(),
input_usage_units: z.number().nullish(),
output_usage_units: z.number().nullish(),
total_usage_units: z.number().nullish(),
input_cost: z.number().nullish(),
output_cost: z.number().nullish(),
total_cost: z.number().nullish(),
provided_input_usage_units: z.number().nullish(),
provided_output_usage_units: z.number().nullish(),
provided_total_usage_units: z.number().nullish(),
provided_input_cost: z.number().nullish(),
provided_output_cost: z.number().nullish(),
provided_total_cost: z.number().nullish(),
prompt_id: z.string().nullish(),
prompt_name: z.string().nullish(),
prompt_version: z.number().nullish(),
});
export type ObservationRecordBaseType = z.infer<
typeof observationRecordBaseSchema
>;
export const observationRecordReadSchema = observationRecordBaseSchema.extend({
created_at: clickhouseStringDateSchema,
updated_at: clickhouseStringDateSchema,
start_time: clickhouseStringDateSchema,
end_time: clickhouseStringDateSchema.nullish(),
completion_start_time: clickhouseStringDateSchema.nullish(),
});
export type ObservationRecordReadType = z.infer<
typeof observationRecordReadSchema
>;
export const observationRecordInsertSchema = observationRecordBaseSchema.extend(
{
created_at: z.number(),
updated_at: z.number(),
start_time: z.number(),
end_time: z.number().nullish(),
completion_start_time: z.number().nullish(),
}
);
export type ObservationRecordInsertType = z.infer<
typeof observationRecordInsertSchema
>;
export const traceRecordBaseSchema = z.object({
id: z.string(),
name: z.string().nullish(),
user_id: z.string().nullish(),
metadata: z.record(z.string()),
release: z.string().nullish(),
version: z.string().nullish(),
project_id: z.string(),
public: z.boolean(),
bookmarked: z.boolean(),
tags: z.array(z.string()),
input: z.string().nullish(),
output: z.string().nullish(),
session_id: z.string().nullish(),
});
export type TraceRecordBaseType = z.infer<typeof traceRecordBaseSchema>;
export const traceRecordReadSchema = traceRecordBaseSchema.extend({
timestamp: clickhouseStringDateSchema,
created_at: clickhouseStringDateSchema,
updated_at: clickhouseStringDateSchema,
});
export type TraceRecordReadType = z.infer<typeof traceRecordReadSchema>;
export const traceRecordInsertSchema = traceRecordBaseSchema.extend({
timestamp: z.number(),
created_at: z.number(),
updated_at: z.number(),
});
export type TraceRecordInsertType = z.infer<typeof traceRecordInsertSchema>;
export const scoreRecordBaseSchema = z.object({
id: z.string(),
project_id: z.string(),
trace_id: z.string(),
observation_id: z.string().nullish(),
name: z.string().nullish(),
value: z.union([z.number(), z.string()]).nullish(),
source: z.string(),
comment: z.string().nullish(),
author_user_id: z.string().nullish(),
config_id: z.string().nullish(),
data_type: z.enum(["NUMERIC", "CATEGORICAL", "BOOLEAN"]).nullish(),
string_value: z.string().nullish(),
});
export type ScoreRecordBaseType = z.infer<typeof scoreRecordBaseSchema>;
export const scoreRecordReadSchema = scoreRecordBaseSchema.extend({
created_at: clickhouseStringDateSchema,
updated_at: clickhouseStringDateSchema,
timestamp: clickhouseStringDateSchema,
});
export type ScoreRecordReadType = z.infer<typeof scoreRecordReadSchema>;
export const scoreRecordInsertSchema = scoreRecordBaseSchema.extend({
created_at: z.number(),
updated_at: z.number(),
timestamp: z.number(),
});
export type ScoreRecordInsertType = z.infer<typeof scoreRecordInsertSchema>;
export const convertTraceReadToInsert = (
record: TraceRecordReadType
): TraceRecordInsertType => {
return {
...record,
created_at: new Date(record.created_at).getTime(),
updated_at: new Date(record.created_at).getTime(),
timestamp: new Date(record.timestamp).getTime(),
};
};
export const convertObservationReadToInsert = (
record: ObservationRecordReadType
): ObservationRecordInsertType => {
return {
...record,
created_at: new Date(record.created_at).getTime(),
updated_at: new Date(record.created_at).getTime(),
start_time: new Date(record.start_time).getTime(),
end_time: record.end_time ? new Date(record.end_time).getTime() : undefined,
completion_start_time: record.completion_start_time
? new Date(record.completion_start_time).getTime()
: undefined,
};
};
export const convertScoreReadToInsert = (
record: ScoreRecordReadType
): ScoreRecordInsertType => {
return {
...record,
created_at: new Date(record.created_at).getTime(),
updated_at: new Date(record.updated_at).getTime(),
timestamp: new Date(record.timestamp).getTime(),
};
};
+13 -13
View File
@@ -26,7 +26,7 @@ const arrayOperatorReplacements = {
export function tableColumnsToSqlFilterAndPrefix(
filters: FilterState,
tableColumns: ColumnDefinition[],
table: TableNames
table: TableNames,
): Prisma.Sql {
const sql = tableColumnsToSqlFilter(filters, tableColumns, table);
if (sql === Prisma.empty) {
@@ -42,14 +42,14 @@ export function tableColumnsToSqlFilterAndPrefix(
export function tableColumnsToSqlFilter(
filters: FilterState,
tableColumns: ColumnDefinition[],
table: TableNames
table: TableNames,
): Prisma.Sql {
const internalFilters = filters.map((filter) => {
// Get column definition to map column to internal name, e.g. "t.id"
const col = tableColumns.find(
(c) =>
// TODO: Only use id instead of name
c.name === filter.column || c.id === filter.column
c.name === filter.column || c.id === filter.column,
);
if (!col) {
logger.error("Invalid filter column", filter.column);
@@ -71,13 +71,13 @@ export function tableColumnsToSqlFilter(
? Prisma.raw(
arrayOperatorReplacements[
filter.operator as keyof typeof arrayOperatorReplacements
]
],
)
: filter.operator in operatorReplacements
? Prisma.raw(
operatorReplacements[
filter.operator as keyof typeof operatorReplacements
]
],
)
: Prisma.raw(filter.operator); //checked by zod
@@ -97,13 +97,13 @@ export function tableColumnsToSqlFilter(
break;
case "stringOptions":
valuePrisma = Prisma.sql`(${Prisma.join(
filter.value.map((v) => Prisma.sql`${v}`)
filter.value.map((v) => Prisma.sql`${v}`),
)})`;
break;
case "arrayOptions":
valuePrisma = Prisma.sql`ARRAY[${Prisma.join(
filter.value.map((v) => Prisma.sql`${v}`),
", "
", ",
)}] `;
break;
@@ -123,12 +123,12 @@ export function tableColumnsToSqlFilter(
filter.type === "string" || filter.type === "stringObject"
? [
["contains", "does not contain", "ends with"].includes(
filter.operator
filter.operator,
)
? Prisma.raw("'%' || ")
: Prisma.empty,
["contains", "does not contain", "starts with"].includes(
filter.operator
filter.operator,
)
? Prisma.raw(" || '%'")
: Prisma.empty,
@@ -152,7 +152,7 @@ export function tableColumnsToSqlFilter(
const castValueToPostgresTypes = (
column: ColumnDefinition,
table: TableNames
table: TableNames,
) => {
return column.name === "type" &&
(table === "observations" ||
@@ -168,7 +168,7 @@ const dateOperators = filterOperators["datetime"];
export const datetimeFilterToPrismaSql = (
safeColumn: string,
operator: (typeof dateOperators)[number],
value: Date
value: Date,
) => {
if (!dateOperators.includes(operator)) {
throw new Error("Invalid operator: " + operator);
@@ -178,12 +178,12 @@ export const datetimeFilterToPrismaSql = (
}
return Prisma.sql`AND ${Prisma.raw(safeColumn)} ${Prisma.raw(
operator
operator,
)} ${value}::timestamp with time zone at time zone 'UTC'`;
};
export const datetimeFilterToPrisma = (
timestampFilter: z.infer<typeof timeFilter>
timestampFilter: z.infer<typeof timeFilter>,
) => {
const prismaTimestampFilter =
timestampFilter.operator === ">="
+10 -8
View File
@@ -9,18 +9,19 @@ export * from "./llm/fetchLLMCompletion";
export * from "./llm/types";
export * from "./utils/DatabaseReadStream";
export * from "./utils/transforms";
export * from "./clickhouse";
export * from "../server/definitions";
export * from "../server/ingestion/IngestionUtils";
export * from "./clickhouse/client";
export * from "./clickhouse/schemaUtils";
export * from "./clickhouse/schema";
export * from "./repositories/definitions";
export * from "../server/ingestion/types";
export * from "../server/ingestion/model-match";
export * from "./ingestion/modelMatch";
export * from "../server/ingestion/types";
export * from "../server/ingestion/validateAndInflateScore";
export * from "./redis/redis";
export * from "./redis/trace-upsert";
export * from "./redis/batch-export";
export * from "./redis/ingestionFlushQueue";
export * from "./redis/legacy-ingestion";
export * from "./redis/traceUpsert";
export * from "./redis/batchExport";
export * from "./redis/legacyIngestion";
export * from "./redis/ingestionQueue";
export * from "./auth/types";
export * from "./ingestion/legacy/index";
export * from "./queues";
@@ -30,3 +31,4 @@ export * from "./filterToPrisma";
export * from "./instrumentation";
export * from "./logger";
export * from "./queries";
export * from "./repositories";
@@ -1,86 +0,0 @@
import { eventTypes, IngestionEventType } from "./types";
const reservedCharsEscapeMap = [
{ reserved: ":", escape: "|%|" },
{ reserved: "_", escape: "|#|" },
];
export enum ClickhouseEntityType {
Trace = "trace",
Score = "score",
Observation = "observation",
SDK_LOG = "sdk-log",
}
export class IngestionUtils {
public static getBufferKey(flushKey: string): string {
const { projectId, eventType, entityId } =
IngestionUtils.parseFlushKey(flushKey);
const sanitizedEntityId = IngestionUtils.escapeReservedChars(entityId);
return (
"ingestionBuffer:" + `${projectId}_${eventType}_${sanitizedEntityId}`
);
}
public static getFlushKey(params: {
projectId: string;
eventType: ClickhouseEntityType;
entityId: string;
batchTimestamp: string;
}): string {
const { projectId, eventType, entityId, batchTimestamp } = params;
const sanitizedEntityId = IngestionUtils.escapeReservedChars(entityId);
return `${projectId}_${eventType}_${sanitizedEntityId}_${batchTimestamp}`;
}
public static parseFlushKey(projectEntityKey: string) {
const split = projectEntityKey.split("_");
if (split.length < 3) {
throw new Error(
`Invalid project entity key format ${projectEntityKey}, expected 3 or 4 parts`
);
}
const [projectId, eventType, escapedEntityId, batchTimestamp] = split;
const entityId = IngestionUtils.unescapeReservedChars(escapedEntityId);
return { projectId, eventType, entityId, batchTimestamp };
}
private static escapeReservedChars(string: string): string {
return reservedCharsEscapeMap.reduce(
(acc, { reserved, escape }) => acc.replaceAll(reserved, escape),
string
);
}
private static unescapeReservedChars(escapedString: string): string {
return reservedCharsEscapeMap.reduce(
(acc, { reserved, escape }) => acc.replaceAll(escape, reserved),
escapedString
);
}
public static getEventType(event: IngestionEventType): ClickhouseEntityType {
switch (event.type) {
case eventTypes.TRACE_CREATE:
return ClickhouseEntityType.Trace;
case eventTypes.OBSERVATION_CREATE:
case eventTypes.OBSERVATION_UPDATE:
case eventTypes.EVENT_CREATE:
case eventTypes.SPAN_CREATE:
case eventTypes.SPAN_UPDATE:
case eventTypes.GENERATION_CREATE:
case eventTypes.GENERATION_UPDATE:
return ClickhouseEntityType.Observation;
case eventTypes.SCORE_CREATE:
return ClickhouseEntityType.Score;
case eventTypes.SDK_LOG:
return ClickhouseEntityType.SDK_LOG;
}
}
}
@@ -1,7 +1,7 @@
import { v4 } from "uuid";
import { type z } from "zod";
import Decimal from "decimal.js";
import { findModel } from "../model-match";
import { findModel } from "../modelMatch";
import {
ObservationEvent,
eventTypes,
@@ -117,11 +117,11 @@ export class ObservationProcessor implements EventProcessor {
projectId: apiScope.projectId,
model:
"model" in this.event.body
? this.event.body.model ?? undefined
? (this.event.body.model ?? undefined)
: undefined,
unit:
"usage" in this.event.body
? this.event.body.usage?.unit ?? undefined
? (this.event.body.usage?.unit ?? undefined)
: undefined,
startTime: this.event.body.startTime
? new Date(this.event.body.startTime)
@@ -158,10 +158,10 @@ export class ObservationProcessor implements EventProcessor {
const newTotalCount =
"usage" in this.event.body
? this.event.body.usage?.total ??
? (this.event.body.usage?.total ??
(newInputCount != null || newOutputCount != null
? (newInputCount ?? 0) + (newOutputCount ?? 0)
: undefined)
: undefined))
: undefined;
const userProvidedTokenCosts = {
@@ -220,7 +220,15 @@ export class ObservationProcessor implements EventProcessor {
logger.warn("Prompt not found for observation", this.event.body);
}
const observationId = this.event.body.id ?? v4();
const observationId =
this.event.body.id ??
(() => {
const newId = v4();
logger.info(
`observation.id is null. Generating for projectId: ${apiScope.projectId}, id: ${newId}`,
);
return newId;
})();
return {
id: observationId,
@@ -245,7 +253,7 @@ export class ObservationProcessor implements EventProcessor {
model: "model" in this.event.body ? this.event.body.model : undefined,
modelParameters:
"modelParameters" in this.event.body
? this.event.body.modelParameters ?? undefined
? (this.event.body.modelParameters ?? undefined)
: undefined,
input: this.event.body.input ?? undefined,
output: this.event.body.output ?? undefined,
@@ -254,7 +262,7 @@ export class ObservationProcessor implements EventProcessor {
totalTokens: newTotalCount,
unit:
"usage" in this.event.body
? this.event.body.usage?.unit ?? internalModel?.unit
? (this.event.body.usage?.unit ?? internalModel?.unit)
: internalModel?.unit,
level: this.event.body.level ?? undefined,
statusMessage: this.event.body.statusMessage ?? undefined,
@@ -300,7 +308,7 @@ export class ObservationProcessor implements EventProcessor {
model: "model" in this.event.body ? this.event.body.model : undefined,
modelParameters:
"modelParameters" in this.event.body
? this.event.body.modelParameters ?? undefined
? (this.event.body.modelParameters ?? undefined)
: undefined,
input: this.event.body.input ?? undefined,
output: this.event.body.output ?? undefined,
@@ -309,7 +317,7 @@ export class ObservationProcessor implements EventProcessor {
totalTokens: newTotalCount,
unit:
"usage" in this.event.body
? this.event.body.usage?.unit ?? internalModel?.unit
? (this.event.body.usage?.unit ?? internalModel?.unit)
: internalModel?.unit,
level: this.event.body.level ?? undefined,
statusMessage: this.event.body.statusMessage ?? undefined,
@@ -359,7 +367,7 @@ export class ObservationProcessor implements EventProcessor {
text: body.input,
});
} else {
logger.info(
logger.debug(
`No input provided, trying to calculate for id: ${existingObservation?.id}`,
);
const observationInput = await prisma.observation.findFirst({
@@ -385,7 +393,7 @@ export class ObservationProcessor implements EventProcessor {
text: body.output,
});
} else {
logger.info(
logger.debug(
`No output provided, trying to calculate for id: ${existingObservation?.id}`,
);
const observationOutput = await prisma.observation.findFirst({
@@ -448,7 +456,7 @@ export class ObservationProcessor implements EventProcessor {
const finalTotalCost =
tokenCounts.total !== undefined && model?.totalPrice
? model.totalPrice.mul(tokenCounts.total)
: finalInputCost ?? finalOutputCost
: (finalInputCost ?? finalOutputCost)
? new Decimal(finalInputCost ?? 0).add(finalOutputCost ?? 0)
: undefined;
@@ -551,7 +559,15 @@ export class TraceProcessor implements EventProcessor {
this.auth(apiScope);
const internalId = body.id ?? v4();
const internalId =
body.id ??
(() => {
const newId = v4();
logger.info(
`trace.id is null. Generating for projectId: ${apiScope.projectId}, id: ${newId}`,
);
return newId;
})();
logger.debug(
`Trying to create trace, project ${apiScope.projectId}, id: ${internalId}`,
@@ -663,7 +679,15 @@ export class ScoreProcessor implements EventProcessor {
this.auth(apiScope);
const id = body.id ?? v4();
const id =
body.id ??
(() => {
const newId = v4();
logger.info(
`score.id is null. Generating for projectId: ${apiScope.projectId}, id: ${newId}`,
);
return newId;
})();
const existingScore = await prisma.score.findFirst({
where: {
@@ -1,72 +0,0 @@
import { type Redis } from "ioredis";
import { QueueJobs } from "../../queues";
import {
getIngestionFlushQueue,
IngestionFlushQueue,
} from "../../redis/ingestionFlushQueue";
import { IngestionUtils } from "../IngestionUtils";
import { IngestionEventType } from "../types";
import { redis } from "../../redis/redis";
import { env } from "../../../env";
import { logger } from "../../logger";
export async function enqueueIngestionEvents(
projectId: string,
events: IngestionEventType[],
) {
const ingestionFlushQueue = getIngestionFlushQueue();
if (!ingestionFlushQueue) {
throw Error("IngestionFlushQueue not initialized");
}
if (!redis) throw Error("Redis connection not available");
const queuedEventPromises: Promise<void>[] = [];
const batchTimestamp = Date.now().toString();
// Use for loop as TS does not narrow redis type in map function
for (const event of events) {
queuedEventPromises.push(
enqueueSingleIngestionEvent(
projectId,
event,
redis,
ingestionFlushQueue,
batchTimestamp,
),
);
}
await Promise.all(queuedEventPromises);
}
async function enqueueSingleIngestionEvent(
projectId: string,
event: IngestionEventType,
redis: Redis,
ingestionFlushQueue: IngestionFlushQueue,
batchTimestamp: string,
): Promise<void> {
if (!("id" in event.body && event.body.id)) {
logger.warn(
`Received ingestion event without id: ${JSON.stringify(event)}`,
);
return;
}
const flushKey = IngestionUtils.getFlushKey({
entityId: event.body.id,
eventType: IngestionUtils.getEventType(event),
projectId,
batchTimestamp,
});
const bufferKey = IngestionUtils.getBufferKey(flushKey);
const serializedEventData = JSON.stringify({ ...event, projectId });
await redis.lpush(bufferKey, serializedEventData);
await redis.expire(bufferKey, env.LANGFUSE_INGESTION_BUFFER_TTL_SECONDS);
await ingestionFlushQueue.add(QueueJobs.FlushIngestionEntity, null, {
jobId: flushKey,
});
}
@@ -1,18 +1,17 @@
import { env } from "node:process";
import z from "zod";
import { ForbiddenError, UnauthorizedError } from "../../../errors";
import { eventTypes, ingestionApiSchema, ingestionEvent } from "../types";
import { eventTypes, ingestionApiSchema, IngestionEventType } from "../types";
import { getProcessorForEvent } from "./EventProcessor";
import { TraceUpsertEventType } from "../../queues";
import {
convertTraceUpsertEventsToRedisEvents,
getTraceUpsertQueue,
} from "../../redis/trace-upsert";
TraceUpsertQueue,
} from "../../redis/traceUpsert";
import { ApiAccessScope } from "../../auth/types";
import { redis } from "../../redis/redis";
import { backOff } from "exponential-backoff";
import { Model } from "../../..";
import { enqueueIngestionEvents } from "./enqueueIngestionEvents";
import { logger } from "../../logger";
export type BatchResult = {
@@ -78,22 +77,13 @@ export const handleBatch = async (
// Decide how to handle the error: rethrow, continue, or push an error object to results
// For example, push an error object:
errors.push({
error: error,
error,
id: singleEvent.id,
type: singleEvent.type,
});
}
}
if (env.CLICKHOUSE_URL) {
try {
await enqueueIngestionEvents(authCheck.scope.projectId, events);
logger.info(`Added ${events.length} ingestion events to queue`);
} catch (err) {
logger.error("Error adding ingestion events to queue", err);
}
}
return { results, errors };
};
@@ -112,7 +102,7 @@ async function retry<T>(request: () => Promise<T>): Promise<T> {
}
const handleSingleEvent = async (
event: z.infer<typeof ingestionEvent>,
event: IngestionEventType,
apiScope: LegacyIngestionAccessScope,
calculateTokenDelegate: (p: {
model: Model;
@@ -132,11 +122,11 @@ const handleSingleEvent = async (
restEvent = rest;
}
logger.info(
logger.debug(
`handling single event ${event.id} of type ${event.type}: ${JSON.stringify({ body: restEvent })}`,
);
const cleanedEvent = ingestionEvent.parse(cleanEvent(event));
const cleanedEvent = cleanEvent(event) as IngestionEventType;
// Deny access to non-score events if the access level is not "all"
// This is an additional safeguard to auth checks in EventProcessor
@@ -198,9 +188,9 @@ export const addTracesToTraceUpsertQueue = async (
try {
if (env.NEXT_PUBLIC_LANGFUSE_CLOUD_REGION && redis) {
logger.info(`Sending ${traceEvents.length} events to worker via Redis`);
logger.debug(`Sending ${traceEvents.length} events to worker via Redis`);
const queue = getTraceUpsertQueue();
const queue = TraceUpsertQueue.getInstance();
if (!queue) {
logger.error("TraceUpsertQueue not initialized");
return;
@@ -3,7 +3,7 @@ import { z } from "zod";
import { NonEmptyString, jsonSchema } from "../../utils/zod";
import { ModelUsageUnit } from "../../constants";
import { ObservationLevel } from "@prisma/client";
import { ObservationLevel, ScoreSource } from "@prisma/client";
export const Usage = z.object({
input: z.number().int().nullish(),
@@ -163,6 +163,7 @@ const BaseScoreBody = z.object({
traceId: z.string(),
observationId: z.string().nullish(),
comment: z.string().nullish(),
source: z.nativeEnum(ScoreSource).default(ScoreSource.API),
});
/**
@@ -76,7 +76,7 @@ function inflateScoreBody(
const { body, projectId, scoreId, config } = params;
const relevantDataType = config?.dataType ?? body.dataType;
const scoreProps = { ...body, id: scoreId, projectId, source: "API" };
const scoreProps = { source: "API", ...body, id: scoreId, projectId };
if (typeof body.value === "number") {
if (relevantDataType && relevantDataType === ScoreDataType.BOOLEAN) {
@@ -1,5 +1,11 @@
import * as opentelemetry from "@opentelemetry/api";
import * as dd from "dd-trace";
import { env } from "../../env";
import {
CloudWatchClient,
PutMetricDataCommand,
} from "@aws-sdk/client-cloudwatch";
import { logger } from "../logger";
// type CallbackFn<T> = () => T;
@@ -32,7 +38,7 @@ export async function instrumentAsync<T>(
return getTracer(ctx.traceScope ?? callback.name).startActiveSpan(
ctx.name,
{
root: !Boolean(ctx.traceContext) && ctx.rootSpan,
root: !ctx.traceContext && ctx.rootSpan,
kind: ctx.spanKind,
},
activeContext,
@@ -66,7 +72,7 @@ export function instrumentSync<T>(
return getTracer(ctx.traceScope ?? callback.name).startActiveSpan(
ctx.name,
{
root: !Boolean(ctx.traceContext) && ctx.rootSpan,
root: !ctx.traceContext && ctx.rootSpan,
kind: ctx.spanKind,
},
activeContext,
@@ -153,6 +159,36 @@ export const addUserToSpan = (
export const getTracer = (name: string) => opentelemetry.trace.getTracer(name);
const cloudWatchClient = new CloudWatchClient();
const cloudWatchLastSubmitted: Record<string, number> = {};
const sendCloudWatchMetric = (key: string, value: number | undefined) => {
const currentTime = Date.now();
const interval = 30 * 1000;
// Check if the function has been executed in the last 30s for this key
if (
!cloudWatchLastSubmitted[key] ||
currentTime - cloudWatchLastSubmitted[key] >= interval
) {
cloudWatchLastSubmitted[key] = currentTime;
cloudWatchClient
.send(
new PutMetricDataCommand({
Namespace: "Langfuse",
MetricData: [
{
MetricName: key,
Value: value ?? 0,
},
],
}),
)
.catch((error) => {
logger.warn("Failed to send metric to CloudWatch", error);
});
}
};
export const recordGauge = (
stat: string,
value?: number | undefined,
@@ -162,6 +198,9 @@ export const recordGauge = (
}
| undefined,
) => {
if (env.ENABLE_AWS_CLOUDWATCH_METRIC_PUBLISHING === "true") {
sendCloudWatchMetric(stat, value);
}
dd.dogstatsd.gauge(stat, value, tags);
};
@@ -170,6 +209,9 @@ export const recordIncrement = (
value?: number | undefined,
tags?: { [tag: string]: string | number } | undefined,
) => {
if (env.ENABLE_AWS_CLOUDWATCH_METRIC_PUBLISHING === "true") {
sendCloudWatchMetric(stat, value);
}
dd.dogstatsd.increment(stat, value, tags);
};
@@ -178,6 +220,9 @@ export const recordHistogram = (
value?: number | undefined,
tags?: { [tag: string]: string | number } | undefined,
) => {
if (env.ENABLE_AWS_CLOUDWATCH_METRIC_PUBLISHING === "true") {
sendCloudWatchMetric(stat, value);
}
dd.dogstatsd.histogram(stat, value, tags);
};
+4 -2
View File
@@ -49,7 +49,7 @@ export const ZodModelConfig = z.object({
top_p: z.coerce.number().optional(),
});
// NOTE: Update docs page when changing this!
// NOTE: Update docs page when changing this! https://langfuse.com/docs/playground#openai-playground--anthropic-playground
export const openAIModels = [
"gpt-4o",
"gpt-4o-2024-08-06",
@@ -76,11 +76,13 @@ export const openAIModels = [
export type OpenAIModel = (typeof openAIModels)[number];
// NOTE: Update docs page when changing this!
// NOTE: Update docs page when changing this! https://langfuse.com/docs/playground#openai-playground--anthropic-playground
export const anthropicModels = [
"claude-3-5-sonnet-20241022",
"claude-3-5-sonnet-20240620",
"claude-3-opus-20240229",
"claude-3-sonnet-20240229",
"claude-3-5-haiku-20241022",
"claude-3-haiku-20240307",
"claude-2.1",
"claude-2.0",
@@ -0,0 +1,384 @@
import { filterOperators } from "../../../interfaces/filters";
function randomCharacters() {
const chars = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz";
let result = "";
const randomArray = new Uint8Array(5);
crypto.getRandomValues(randomArray);
randomArray.forEach((number) => {
result += chars[number % chars.length];
});
return result;
}
export interface Filter {
apply(): ClickhouseFilter;
clickhouseTable: string;
operator: (typeof filterOperators)[keyof typeof filterOperators][number];
field: string;
}
type ClickhouseFilter = {
query: string;
params: { [x: string]: any } | {};
};
export class StringFilter implements Filter {
public clickhouseTable: string;
public field: string;
public value: string;
public operator: (typeof filterOperators)["string"][number];
protected tablePrefix?: string;
constructor(opts: {
clickhouseTable: string;
field: string;
operator: (typeof filterOperators)["string"][number];
value: string;
tablePrefix?: string;
}) {
this.clickhouseTable = opts.clickhouseTable;
this.field = opts.field;
this.value = opts.value;
this.operator = opts.operator;
this.tablePrefix = opts.tablePrefix;
}
apply(): ClickhouseFilter {
const varName = `stringFilter${randomCharacters()}`;
const fieldWithPrefix = `${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field}`;
let query: string;
switch (this.operator) {
case "=":
query = `${fieldWithPrefix} = {${varName}: String}`;
break;
case "contains":
query = `position(${fieldWithPrefix}, {${varName}: String}) > 0`;
break;
case "does not contain":
query = `position(${fieldWithPrefix}, {${varName}: String}) = 0`;
break;
case "starts with":
query = `startsWith(${fieldWithPrefix}, {${varName}: String})`;
break;
case "ends with":
query = `endsWith(${fieldWithPrefix}, {${varName}: String})`;
break;
default:
throw new Error(`Unsupported operator: ${this.operator}`);
}
return {
query: query,
params: { [varName]: this.value },
};
}
}
export class NumberFilter implements Filter {
public clickhouseTable: string;
public field: string;
public value: number;
public operator: (typeof filterOperators)["number"][number];
protected tablePrefix?: string;
constructor(opts: {
clickhouseTable: string;
field: string;
operator: (typeof filterOperators)["number"][number];
value: number;
tablePrefix?: string;
}) {
this.clickhouseTable = opts.clickhouseTable;
this.field = opts.field;
this.value = opts.value;
this.operator = opts.operator;
this.tablePrefix = opts.tablePrefix;
}
apply(): ClickhouseFilter {
const uid = randomCharacters();
const varName = `numberFilter${uid}`;
return {
query: `${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field} ${this.operator} {${varName}: Decimal}`,
params: { [varName]: this.value },
};
}
}
export class DateTimeFilter implements Filter {
public clickhouseTable: string;
public field: string;
public value: Date;
public operator: (typeof filterOperators)["datetime"][number];
protected tablePrefix?: string;
constructor(opts: {
clickhouseTable: string;
field: string;
operator: (typeof filterOperators)["datetime"][number];
value: Date;
tablePrefix?: string;
}) {
this.clickhouseTable = opts.clickhouseTable;
this.field = opts.field;
this.value = opts.value;
this.operator = opts.operator;
this.tablePrefix = opts.tablePrefix;
}
apply(): ClickhouseFilter {
const uid = randomCharacters();
const varName = `dateTimeFilter${uid}`;
return {
query: `${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field} ${this.operator} {${varName}: DateTime64(3)}`,
params: { [varName]: new Date(this.value).getTime() },
};
}
}
export class StringOptionsFilter implements Filter {
public clickhouseTable: string;
public field: string;
public values: string[];
public operator: (typeof filterOperators.stringOptions)[number];
protected tablePrefix?: string;
constructor(opts: {
clickhouseTable: string;
field: string;
operator: (typeof filterOperators.stringOptions)[number];
values: string[];
tablePrefix?: string;
}) {
this.clickhouseTable = opts.clickhouseTable;
this.field = opts.field;
this.values = opts.values;
this.operator = opts.operator;
this.tablePrefix = opts.tablePrefix;
}
apply(): ClickhouseFilter {
const uid = randomCharacters();
const varName = `stringOptionsFilter${uid}`;
return {
query:
this.operator === "any of"
? `has({${varName}: Array(String)}, ${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field}) = True`
: `has({${varName}: Array(String)}, ${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field}) = False`,
params: { [varName]: this.values },
};
}
}
// stringObject filter is used when we want to filter on a key value pair in a clickhouse map.
// As we use the MAP form clickhouse, we can only filter efficiently on the first level of a json obj.
export class StringObjectFilter implements Filter {
public clickhouseTable: string;
public field: string;
public key: string;
public value: string;
public operator: (typeof filterOperators)["stringObject"][number];
protected tablePrefix?: string;
constructor(opts: {
clickhouseTable: string;
field: string;
operator: (typeof filterOperators)["stringObject"][number];
key: string;
value: string;
tablePrefix?: string;
}) {
this.clickhouseTable = opts.clickhouseTable;
this.field = opts.field;
this.value = opts.value;
this.operator = opts.operator;
this.tablePrefix = opts.tablePrefix;
this.key = opts.key;
}
apply(): ClickhouseFilter {
const varKeyName = `stringObjectKeyFilter${randomCharacters()}`;
const varValueName = `stringObjectValueFilter${randomCharacters()}`;
const column = `${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field}`;
// const query: `${column}['{varKeyName: String}'] ${this.operator} {${varValueName}: String}`,
let query: string;
switch (this.operator) {
case "=":
query = `${column}[{${varKeyName}: String}] = {${varValueName}: String}`;
break;
case "contains":
query = `position(${column}[{${varKeyName}: String}], {${varValueName}: String}) > 0`;
break;
case "does not contain":
query = `position(${column}[{${varKeyName}: String}], {${varValueName}: String}) = 0`;
break;
case "starts with":
query = `startsWith(${column}[{${varKeyName}: String}], {${varValueName}: String})`;
break;
case "ends with":
query = `endsWith(${column}[{${varKeyName}: String}], {${varValueName}: String})`;
break;
default:
throw new Error(`Unsupported operator: ${this.operator}`);
}
return {
query,
params: { [varKeyName]: this.key, [varValueName]: this.value },
};
}
}
// this is used when we want to filter multiple values on a clickhouse column which is also an array
export class ArrayOptionsFilter implements Filter {
public clickhouseTable: string;
public field: string;
public values: string[];
public operator: (typeof filterOperators.arrayOptions)[number];
protected tablePrefix?: string;
constructor(opts: {
clickhouseTable: string;
field: string;
operator: (typeof filterOperators.arrayOptions)[number];
values: string[];
tablePrefix?: string;
}) {
this.clickhouseTable = opts.clickhouseTable;
this.field = opts.field;
this.values = opts.values;
this.operator = opts.operator;
this.tablePrefix = opts.tablePrefix;
}
apply(): ClickhouseFilter {
const uid = randomCharacters();
const varName = `arrayOptionsFilter${uid}`;
let query: string;
switch (this.operator) {
case "any of":
query = `hasAny({${varName}: Array(String)}, ${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field}) = True`;
break;
case "none of":
query = `hasAny({${varName}: Array(String)}, ${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field}) = False`;
break;
case "all of":
query = `arrayAll(x -> has({${varName}: Array(String)}, x), ${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field}) = True`;
break;
default:
throw new Error(`Unsupported operator: ${this.operator}`);
}
return {
query,
params: { [varName]: this.values },
};
}
}
export class NumberObjectFilter implements Filter {
public clickhouseTable: string;
public field: string;
public key: string;
public value: number;
public operator: (typeof filterOperators)["numberObject"][number];
protected tablePrefix?: string;
constructor(opts: {
clickhouseTable: string;
field: string;
operator: (typeof filterOperators)["numberObject"][number];
key: string;
value: number;
tablePrefix?: string;
}) {
this.clickhouseTable = opts.clickhouseTable;
this.field = opts.field;
this.value = opts.value;
this.operator = opts.operator;
this.tablePrefix = opts.tablePrefix;
this.key = opts.key;
}
apply(): ClickhouseFilter {
const varKeyName = `numberObjectKeyFilter${randomCharacters()}`;
const varValueName = `numberObjectValueFilter${randomCharacters()}`;
const column = `${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field}`;
return {
query: `empty(arrayFilter(x -> (((x.1) = {${varKeyName}: String}) AND ((x.2) ${this.operator} {${varValueName}: Decimal})), ${column})) = 0`,
params: { [varKeyName]: this.key, [varValueName]: this.value },
};
}
}
export class BooleanFilter implements Filter {
public clickhouseTable: string;
public field: string;
public operator: (typeof filterOperators)["boolean"][number];
public value: boolean;
protected tablePrefix?: string;
constructor(opts: {
clickhouseTable: string;
field: string;
operator: (typeof filterOperators)["boolean"][number];
value: boolean;
tablePrefix?: string;
}) {
this.clickhouseTable = opts.clickhouseTable;
this.field = opts.field;
this.value = opts.value;
this.tablePrefix = opts.tablePrefix;
this.operator = opts.operator;
}
apply(): ClickhouseFilter {
const uid = randomCharacters();
const varName = `booleanFilter${uid}`;
return {
query: `${this.tablePrefix ? this.tablePrefix + "." : ""}${this.field} ${this.operator} {${varName}: Boolean}`,
params: { [varName]: this.value },
};
}
}
export class FilterList {
private filters: Filter[];
constructor(filters: Filter[]) {
this.filters = filters;
}
push(...filter: Filter[]) {
this.filters.push(...filter);
}
find(predicate: (filter: Filter) => boolean) {
return this.filters.find(predicate);
}
public apply(): ClickhouseFilter {
if (this.filters.length === 0) {
return {
query: "",
params: {},
};
}
const compiledQueries = this.filters.map((filter) => filter.apply());
const { params, queries } = compiledQueries.reduce(
(acc, { params, query }) => {
acc.params = { ...acc.params, ...params };
acc.queries.push(query);
return acc;
},
{ params: {}, queries: [] as string[] },
);
return {
query: queries.join(" AND "),
params,
};
}
}
@@ -0,0 +1,173 @@
import z from "zod";
import { singleFilter } from "../../../interfaces/filters";
import { FilterCondition } from "../../../types";
import { isValidTableName } from "../../clickhouse/schemaUtils";
import { logger } from "../../logger";
import { UiColumnMapping } from "../../../tableDefinitions";
import {
StringFilter,
DateTimeFilter,
StringOptionsFilter,
FilterList,
NumberFilter,
ArrayOptionsFilter,
BooleanFilter,
NumberObjectFilter,
StringObjectFilter,
} from "./clickhouse-filter";
export class QueryBuilderError extends Error {
constructor(message: string) {
super(message);
this.name = "QueryBuilderError";
}
}
// This function ensures that the user only selects valid columns from the clickhouse schema.
// The filter property in this column needs to be zod verified.
// User input for values (e.g. project_id = <value>) are sent to Clickhouse as parameters to prevent SQL injection
export const createFilterFromFilterState = (
filter: FilterCondition[],
columnMapping: UiColumnMapping[],
) => {
return filter.map((frontEndFilter) => {
// checks if the column exists in the clickhouse schema
const column = matchAndVerifyTracesUiColumn(frontEndFilter, columnMapping);
switch (frontEndFilter.type) {
case "string":
return new StringFilter({
clickhouseTable: column.clickhouseTableName,
field: column.clickhouseSelect,
operator: frontEndFilter.operator,
value: frontEndFilter.value,
tablePrefix: column.queryPrefix,
});
case "datetime":
return new DateTimeFilter({
clickhouseTable: column.clickhouseTableName,
field: column.clickhouseSelect,
operator: frontEndFilter.operator,
value: frontEndFilter.value,
tablePrefix: column.queryPrefix,
});
case "stringOptions":
return new StringOptionsFilter({
clickhouseTable: column.clickhouseTableName,
field: column.clickhouseSelect,
operator: frontEndFilter.operator,
values: frontEndFilter.value,
tablePrefix: column.queryPrefix,
});
case "number":
return new NumberFilter({
clickhouseTable: column.clickhouseTableName,
field: column.clickhouseSelect,
operator: frontEndFilter.operator,
value: frontEndFilter.value,
tablePrefix: column.queryPrefix,
});
case "arrayOptions":
return new ArrayOptionsFilter({
clickhouseTable: column.clickhouseTableName,
field: column.clickhouseSelect,
operator: frontEndFilter.operator,
values: frontEndFilter.value,
tablePrefix: column.queryPrefix,
});
case "boolean":
return new BooleanFilter({
clickhouseTable: column.clickhouseTableName,
field: column.clickhouseSelect,
value: frontEndFilter.value,
operator: frontEndFilter.operator,
tablePrefix: column.queryPrefix,
});
case "numberObject":
return new NumberObjectFilter({
clickhouseTable: column.clickhouseTableName,
field: column.clickhouseSelect,
key: frontEndFilter.key,
operator: frontEndFilter.operator,
value: frontEndFilter.value,
tablePrefix: column.queryPrefix,
});
case "stringObject":
return new StringObjectFilter({
clickhouseTable: column.clickhouseTableName,
field: column.clickhouseSelect,
operator: frontEndFilter.operator,
key: frontEndFilter.key,
value: frontEndFilter.value,
tablePrefix: column.queryPrefix,
});
default:
logger.error(`Invalid filter type: ${JSON.stringify(frontEndFilter)}`);
throw new QueryBuilderError(`Invalid filter type`);
}
});
};
const matchAndVerifyTracesUiColumn = (
filter: z.infer<typeof singleFilter>,
uiTableDefinitions: UiColumnMapping[],
) => {
// tries to match the column name to the clickhouse table name
logger.debug(`Filter to match: ${JSON.stringify(filter)}`);
const uiTable = uiTableDefinitions.find(
(col) =>
col.uiTableName === filter.column || col.uiTableId === filter.column, // matches on the NAME of the column in the UI.
);
if (!uiTable) {
throw new QueryBuilderError(
`Column ${filter.column} does not match a UI / CH table mapping.`,
);
}
if (!isValidTableName(uiTable.clickhouseTableName)) {
throw new QueryBuilderError(
`Invalid clickhouse table name: ${uiTable.clickhouseTableName}`,
);
}
return uiTable;
};
export function getProjectIdDefaultFilter(
projectId: string,
opts: { tracesPrefix: string },
): {
tracesFilter: FilterList;
scoresFilter: FilterList;
observationsFilter: FilterList;
} {
return {
tracesFilter: new FilterList([
new StringFilter({
clickhouseTable: "traces",
field: "project_id",
operator: "=",
value: projectId,
tablePrefix: opts.tracesPrefix,
}),
]),
scoresFilter: new FilterList([
new StringFilter({
clickhouseTable: "scores",
field: "project_id",
operator: "=",
value: projectId,
}),
]),
observationsFilter: new FilterList([
new StringFilter({
clickhouseTable: "observations",
field: "project_id",
operator: "=",
value: projectId,
}),
]),
};
}
@@ -0,0 +1,33 @@
import z from "zod";
import { OrderByState } from "../../../interfaces/orderBy";
import { UiColumnMapping } from "../../../tableDefinitions";
import { logger } from "../../logger";
export function orderByToClickhouseSql(
orderBy: OrderByState,
tableColumns: UiColumnMapping[],
): string {
if (!orderBy) {
return "";
}
// Get column definition to map column to internal name, e.g. "t.id"
const col = tableColumns.find(
(c) => c.uiTableName === orderBy.column || c.uiTableId === orderBy.column,
);
if (!col) {
logger.warn("Invalid order by column", orderBy.column);
throw new Error("Invalid order by column: " + orderBy.column);
}
// Assert that orderBy.order is either "asc" or "desc"
const orderByOrder = z.enum(["ASC", "DESC"]);
const order = orderByOrder.safeParse(orderBy.order);
if (!order.success) {
logger.warn("Invalid order", orderBy.order);
throw new Error("Invalid order: " + orderBy.order);
}
// Both column and order are safe, can use raw SQL
return `ORDER BY ${col.queryPrefix ? col.queryPrefix + "." : ""}${col.clickhouseSelect} ${order.data}`;
}
@@ -0,0 +1,20 @@
const regexIndefiniteCharacters = "%";
export const clickhouseSearchCondition = (query?: string) => {
return {
query: query
? `
AND (
id ILIKE {searchString: String} OR
user_id ILIKE {searchString: String} OR
name ILIKE {searchString: String}
)
`
: "",
params: query
? {
searchString: `${regexIndefiniteCharacters}${query}${regexIndefiniteCharacters}`,
}
: {},
};
};
@@ -9,12 +9,10 @@ import { TableFilters } from "./types";
type AdditionalObservationFields = {
traceName: string | null;
promptName: string | null;
promptVersion: string | null;
traceTags: Array<string>;
};
type FullObservation = AdditionalObservationFields & ObservationView;
export type FullObservation = AdditionalObservationFields & ObservationView;
export type FullObservations = Array<FullObservation>;
@@ -40,17 +38,17 @@ export function parseGetAllGenerationsInput(filters: TableFilters) {
const filterCondition = tableColumnsToSqlFilterAndPrefix(
filters.filter ?? [],
observationsTableCols,
"observations"
"observations",
);
const orderByCondition = orderByToPrismaSql(
filters.orderBy,
observationsTableCols
observationsTableCols,
);
// to improve query performance, add timeseries filter to observation queries as well
const startTimeFilter = filters.filter?.find(
(f) => f.column === "Start Time" && f.type === "datetime"
(f) => f.column === "Start Time" && f.type === "datetime",
);
const datetimeFilter =
@@ -58,7 +56,7 @@ export function parseGetAllGenerationsInput(filters: TableFilters) {
? datetimeFilterToPrismaSql(
"start_time",
startTimeFilter.operator,
startTimeFilter.value
startTimeFilter.value,
)
: Prisma.empty;
@@ -4,7 +4,7 @@ import {
datetimeFilterToPrismaSql,
tableColumnsToSqlFilterAndPrefix,
} from "../filterToPrisma";
import { tracesTableCols } from "../../tracesTable";
import { tracesTableCols } from "../../tableDefinitions/tracesTable";
import { orderByToPrismaSql } from "../orderByToPrisma";
export function parseTraceAllFilters(input: TableFilters) {
+26 -10
View File
@@ -1,5 +1,5 @@
import { z } from "zod";
import { eventTypes, ingestionBatchEvent, TCarrier } from ".";
import { eventTypes, ingestionBatchEvent } from ".";
export enum EventName {
TraceUpsert = "TraceUpsert",
@@ -28,7 +28,7 @@ export const LegacyIngestionEventMeta = z.object({
type: z.nativeEnum(eventTypes),
eventBodyId: z.string(),
eventId: z.string(),
})
}),
),
authCheck: z.object({
validKey: z.literal(true),
@@ -44,6 +44,20 @@ export const LegacyIngestionEvent = z.discriminatedUnion("useS3EventStore", [
LegacyIngestionEventMeta,
]);
export const IngestionEvent = z.object({
data: z.object({
type: z.nativeEnum(eventTypes),
eventBodyId: z.string(),
}),
authCheck: z.object({
validKey: z.literal(true),
scope: z.object({
projectId: z.string(),
accessLevel: z.enum(["all", "scores"]),
}),
}),
});
export const BatchExportJobSchema = z.object({
projectId: z.string(),
batchExportId: z.string(),
@@ -62,6 +76,7 @@ export type BatchExportJobType = z.infer<typeof BatchExportJobSchema>;
export type TraceUpsertEventType = z.infer<typeof TraceUpsertEventSchema>;
export type EvalExecutionEventType = z.infer<typeof EvalExecutionEvent>;
export type LegacyIngestionEventType = z.infer<typeof LegacyIngestionEvent>;
export type IngestionEventQueueType = z.infer<typeof IngestionEvent>;
export const EventBodySchema = z.union([
z.object({
@@ -83,9 +98,8 @@ export enum QueueName {
TraceUpsert = "trace-upsert", // Ingestion pipeline adds events on each Trace upsert
EvaluationExecution = "evaluation-execution-queue", // Worker executes Evals
BatchExport = "batch-export-queue",
RepeatQueue = "repeat-queue",
IngestionFlushQueue = "ingestion-flush-queue",
LegacyIngestionQueue = "legacy-ingestion-queue",
IngestionQueue = "ingestion-queue", // Process single events with S3-merge
LegacyIngestionQueue = "legacy-ingestion-queue", // Used for batch processing of Ingestion
CloudUsageMeteringQueue = "cloud-usage-metering-queue",
}
@@ -94,9 +108,9 @@ export enum QueueJobs {
EvaluationExecution = "evaluation-execution-job",
BatchExportJob = "batch-export-job",
EnqueueBatchExportJobs = "enqueue-batch-export-jobs",
FlushIngestionEntity = "flush-ingestion-entity",
LegacyIngestionJob = "legacy-ingestion-job",
CloudUsageMeteringJob = "cloud-usage-metering-job",
IngestionJob = "ingestion-job",
}
export type TQueueJobTypes = {
@@ -105,27 +119,29 @@ export type TQueueJobTypes = {
id: string;
payload: TraceUpsertEventType;
name: QueueJobs.TraceUpsert;
_tracecontext?: TCarrier;
};
[QueueName.EvaluationExecution]: {
timestamp: Date;
id: string;
payload: EvalExecutionEventType;
name: QueueJobs.EvaluationExecution;
_tracecontext?: TCarrier;
};
[QueueName.BatchExport]: {
timestamp: Date;
id: string;
payload: BatchExportJobType;
name: QueueJobs.BatchExportJob;
_tracecontext?: TCarrier;
};
[QueueName.LegacyIngestionQueue]: {
timestamp: Date;
id: string;
payload: LegacyIngestionEventType;
name: QueueJobs.LegacyIngestionJob;
_tracecontext?: TCarrier;
};
[QueueName.IngestionQueue]: {
timestamp: Date;
id: string;
payload: IngestionEventQueueType;
name: QueueJobs.IngestionJob;
};
};
@@ -1,29 +0,0 @@
import { Queue } from "bullmq";
import { QueueName, TQueueJobTypes } from "../queues";
import { createNewRedisInstance } from "./redis";
let batchExportQueue: Queue<TQueueJobTypes[QueueName.BatchExport]> | null =
null;
export const getBatchExportQueue = () => {
if (batchExportQueue) return batchExportQueue;
const connection = createNewRedisInstance();
batchExportQueue = connection
? new Queue<TQueueJobTypes[QueueName.BatchExport]>(QueueName.BatchExport, {
connection: connection,
defaultJobOptions: {
removeOnComplete: true,
removeOnFail: 10_000,
attempts: 2,
backoff: {
type: "exponential",
delay: 5000,
},
},
})
: null;
return batchExportQueue;
};
@@ -0,0 +1,44 @@
import { Queue } from "bullmq";
import { QueueName, TQueueJobTypes } from "../queues";
import { createNewRedisInstance, redisQueueRetryOptions } from "./redis";
import { logger } from "../logger";
export class BatchExportQueue {
private static instance: Queue<TQueueJobTypes[QueueName.BatchExport]> | null =
null;
public static getInstance(): Queue<
TQueueJobTypes[QueueName.BatchExport]
> | null {
if (BatchExportQueue.instance) return BatchExportQueue.instance;
const newRedis = createNewRedisInstance({
enableOfflineQueue: false,
...redisQueueRetryOptions,
});
BatchExportQueue.instance = newRedis
? new Queue<TQueueJobTypes[QueueName.BatchExport]>(
QueueName.BatchExport,
{
connection: newRedis,
defaultJobOptions: {
removeOnComplete: true,
removeOnFail: 10_000,
attempts: 2,
backoff: {
type: "exponential",
delay: 5000,
},
},
},
)
: null;
BatchExportQueue.instance?.on("error", (err) => {
logger.error("BatchExportQueue error", err);
});
return BatchExportQueue.instance;
}
}
@@ -1,33 +0,0 @@
import { Queue } from "bullmq";
import { env } from "../../env";
import { createNewRedisInstance } from "../redis/redis";
import { QueueName } from "../queues";
export type IngestionFlushQueue = Queue<null>;
let ingestionFlushQueue: IngestionFlushQueue | null = null;
export const getIngestionFlushQueue = () => {
if (ingestionFlushQueue) return ingestionFlushQueue;
const connection = createNewRedisInstance();
ingestionFlushQueue = connection
? new Queue<null>(QueueName.IngestionFlushQueue, {
connection: connection,
defaultJobOptions: {
removeOnComplete: true, // Important: If not true, new jobs for that ID would be ignored as jobs in the complete set are still considered as part of the queue
removeOnFail: 100_000,
delay: env.LANGFUSE_INGESTION_FLUSH_DELAY_MS,
attempts: env.LANGFUSE_INGESTION_FLUSH_ATTEMPTS,
backoff: {
type: "exponential",
delay: 5000,
},
},
})
: null;
return ingestionFlushQueue;
};
@@ -0,0 +1,45 @@
import { Queue } from "bullmq";
import { QueueName, TQueueJobTypes } from "../queues";
import { createNewRedisInstance, redisQueueRetryOptions } from "./redis";
import { logger } from "../logger";
export class IngestionQueue {
private static instance: Queue<
TQueueJobTypes[QueueName.IngestionQueue]
> | null = null;
public static getInstance(): Queue<
TQueueJobTypes[QueueName.IngestionQueue]
> | null {
if (IngestionQueue.instance) return IngestionQueue.instance;
const newRedis = createNewRedisInstance({
enableOfflineQueue: false,
...redisQueueRetryOptions,
});
IngestionQueue.instance = newRedis
? new Queue<TQueueJobTypes[QueueName.IngestionQueue]>(
QueueName.IngestionQueue,
{
connection: newRedis,
defaultJobOptions: {
removeOnComplete: true,
removeOnFail: 100_000,
attempts: 5,
backoff: {
type: "exponential",
delay: 5000,
},
},
},
)
: null;
IngestionQueue.instance?.on("error", (err) => {
logger.error("IngestionQueue error", err);
});
return IngestionQueue.instance;
}
}
@@ -1,6 +1,7 @@
import { Queue } from "bullmq";
import { QueueName, TQueueJobTypes } from "../queues";
import { createNewRedisInstance } from "./redis";
import { createNewRedisInstance, redisQueueRetryOptions } from "./redis";
import { logger } from "../logger";
export class LegacyIngestionQueue {
private static instance: Queue<
@@ -12,7 +13,10 @@ export class LegacyIngestionQueue {
> | null {
if (LegacyIngestionQueue.instance) return LegacyIngestionQueue.instance;
const newRedis = createNewRedisInstance({ enableOfflineQueue: false });
const newRedis = createNewRedisInstance({
enableOfflineQueue: false,
...redisQueueRetryOptions,
});
LegacyIngestionQueue.instance = newRedis
? new Queue<TQueueJobTypes[QueueName.LegacyIngestionQueue]>(
@@ -21,7 +25,7 @@ export class LegacyIngestionQueue {
connection: newRedis,
defaultJobOptions: {
removeOnComplete: true,
removeOnFail: 100_000,
removeOnFail: 500_000,
attempts: 5,
backoff: {
type: "exponential",
@@ -32,6 +36,10 @@ export class LegacyIngestionQueue {
)
: null;
LegacyIngestionQueue.instance?.on("error", (err) => {
logger.error("LegacyIngestionQueue error", err);
});
return LegacyIngestionQueue.instance;
}
}
+28 -5
View File
@@ -2,13 +2,30 @@ import Redis, { RedisOptions } from "ioredis";
import { env } from "../../env";
import { logger } from "../logger";
const defaultRedisOptions: Partial<RedisOptions> = {
maxRetriesPerRequest: null,
enableAutoPipelining: env.REDIS_ENABLE_AUTO_PIPELINING === "true",
};
export const redisQueueRetryOptions: Partial<RedisOptions> = {
retryStrategy: (times: number) => {
// Retries forever. Waits at least 1s and at most 20s between retries.
logger.warn(`Connection to redis lost. Retry attempt: ${times}`);
return Math.max(Math.min(Math.exp(times), 20000), 1000);
},
reconnectOnError: (err) => {
// Reconnects on READONLY errors and auto-retries the command.
logger.warn(`Redis connection error: ${err.message}`);
return err.message.includes("READONLY") ? 2 : false;
},
};
export const createNewRedisInstance = (
additionalOptions: Partial<RedisOptions> = {},
) => {
return env.REDIS_CONNECTION_STRING
const instance = env.REDIS_CONNECTION_STRING
? new Redis(env.REDIS_CONNECTION_STRING, {
maxRetriesPerRequest: null,
enableAutoPipelining: env.REDIS_ENABLE_AUTO_PIPELINING === "true",
...defaultRedisOptions,
...additionalOptions,
})
: env.REDIS_HOST
@@ -16,11 +33,16 @@ export const createNewRedisInstance = (
host: String(env.REDIS_HOST),
port: Number(env.REDIS_PORT),
password: String(env.REDIS_AUTH),
maxRetriesPerRequest: null, // Set to `null` to disable retrying
enableAutoPipelining: env.REDIS_ENABLE_AUTO_PIPELINING === "true",
...defaultRedisOptions,
...additionalOptions,
})
: null;
instance?.on("error", (error) => {
logger.error("Redis error", error);
});
return instance;
};
const createRedisClient = () => {
@@ -31,6 +53,7 @@ const createRedisClient = () => {
return null;
}
};
declare global {
// eslint-disable-next-line no-var
var redis: undefined | ReturnType<typeof createRedisClient>;
@@ -1,76 +0,0 @@
import { randomUUID } from "crypto";
import {
QueueJobs,
QueueName,
TQueueJobTypes,
TraceUpsertEventType,
} from "../queues";
import { Queue } from "bullmq";
import { createNewRedisInstance } from "./redis";
let traceUpsertQueue: Queue<TQueueJobTypes[QueueName.TraceUpsert]> | null =
null;
export const getTraceUpsertQueue = () => {
if (traceUpsertQueue) return traceUpsertQueue;
const connection = createNewRedisInstance();
traceUpsertQueue = connection
? new Queue<TQueueJobTypes[QueueName.TraceUpsert]>(QueueName.TraceUpsert, {
connection: connection,
defaultJobOptions: {
removeOnComplete: 100, // Important: If not true, new jobs for that ID would be ignored as jobs in the complete set are still considered as part of the queue
removeOnFail: 100_000,
attempts: 5,
backoff: {
type: "exponential",
delay: 5000,
},
},
})
: null;
return traceUpsertQueue;
};
export function convertTraceUpsertEventsToRedisEvents(
events: TraceUpsertEventType[],
) {
const uniqueTracesPerProject = events.reduce((acc, event) => {
if (!acc.get(event.projectId)) {
acc.set(event.projectId, new Set());
}
acc.get(event.projectId)?.add(event.traceId);
return acc;
}, new Map<string, Set<string>>());
const jobs = [...uniqueTracesPerProject.entries()]
.map((tracesPerProject) => {
const [projectId, traceIds] = tracesPerProject;
return [...traceIds].map((traceId) => ({
name: QueueJobs.TraceUpsert,
data: {
payload: {
projectId,
traceId,
},
id: randomUUID(),
timestamp: new Date(),
name: QueueJobs.TraceUpsert as const,
},
opts: {
removeOnFail: 1_000,
removeOnComplete: true,
attempts: 5,
backoff: {
type: "exponential",
delay: 1000,
},
},
}));
})
.flat();
return jobs;
}
@@ -0,0 +1,90 @@
import { randomUUID } from "crypto";
import {
QueueJobs,
QueueName,
TQueueJobTypes,
TraceUpsertEventType,
} from "../queues";
import { Queue } from "bullmq";
import { createNewRedisInstance, redisQueueRetryOptions } from "./redis";
import { logger } from "../logger";
export class TraceUpsertQueue {
private static instance: Queue<TQueueJobTypes[QueueName.TraceUpsert]> | null =
null;
public static getInstance(): Queue<
TQueueJobTypes[QueueName.TraceUpsert]
> | null {
if (TraceUpsertQueue.instance) return TraceUpsertQueue.instance;
const newRedis = createNewRedisInstance({
enableOfflineQueue: false,
...redisQueueRetryOptions,
});
TraceUpsertQueue.instance = newRedis
? new Queue<TQueueJobTypes[QueueName.TraceUpsert]>(
QueueName.TraceUpsert,
{
connection: newRedis,
defaultJobOptions: {
removeOnComplete: 100, // Important: If not true, new jobs for that ID would be ignored as jobs in the complete set are still considered as part of the queue
removeOnFail: 100_000,
attempts: 5,
backoff: {
type: "exponential",
delay: 5000,
},
},
},
)
: null;
TraceUpsertQueue.instance?.on("error", (err) => {
logger.error("TraceUpsertQueue error", err);
});
return TraceUpsertQueue.instance;
}
}
export function convertTraceUpsertEventsToRedisEvents(
events: TraceUpsertEventType[],
) {
const uniqueTracesPerProject = events.reduce((acc, event) => {
if (!acc.get(event.projectId)) {
acc.set(event.projectId, new Set());
}
acc.get(event.projectId)?.add(event.traceId);
return acc;
}, new Map<string, Set<string>>());
return [...uniqueTracesPerProject.entries()]
.map((tracesPerProject) => {
const [projectId, traceIds] = tracesPerProject;
return [...traceIds].map((traceId) => ({
name: QueueJobs.TraceUpsert,
data: {
payload: {
projectId,
traceId,
},
id: randomUUID(),
timestamp: new Date(),
name: QueueJobs.TraceUpsert as const,
},
opts: {
removeOnFail: 1_000,
removeOnComplete: true,
attempts: 5,
backoff: {
type: "exponential",
delay: 1000,
},
},
}));
})
.flat();
}
@@ -0,0 +1,15 @@
## Repository docs
### Guarantees for relating data within Langfuse
- Finding a Trace based on an observation [Linear](https://linear.app/langfuse/issue/LFE-2745/improve-generations-table-query-performance)
- Traces can occur earlier than an observation.
- 96% of observations.start_time occur 2 mins later than the trace.timestamp
- There is a very large long-tail. Hence we will use a 2-day (2880 min) look back for now.
- Finding an Observation based on a Trace [Linear](https://linear.app/langfuse/issue/LFE-2409/table-queries)
- Observations have a very high likelihood of happening after the trace.
- 97% of traces.timestamp occur 2 mins earlier than the observation.start_time
- We will maintain a 1 hour cutoff for now.
- Finding traces/observations based on a Score timestamp
- For scores we have a very high likelihood of happening after the trace / observation.
- We maintain a 1 hour cutoff for now.
@@ -0,0 +1,186 @@
import { env } from "../../env";
import {
clickhouseClient,
convertDateToClickhouseDateTime,
} from "../clickhouse/client";
import { logger } from "../logger";
import { instrumentAsync } from "../instrumentation";
import { S3StorageService } from "../services/S3StorageService";
import { randomUUID } from "crypto";
import { getClickhouseEntityType } from "../clickhouse/schemaUtils";
let s3StorageServiceClient: S3StorageService;
const getS3StorageServiceClient = (bucketName: string): S3StorageService => {
if (!s3StorageServiceClient) {
s3StorageServiceClient = new S3StorageService({
bucketName,
accessKeyId: env.LANGFUSE_S3_EVENT_UPLOAD_ACCESS_KEY_ID,
secretAccessKey: env.LANGFUSE_S3_EVENT_UPLOAD_SECRET_ACCESS_KEY,
endpoint: env.LANGFUSE_S3_EVENT_UPLOAD_ENDPOINT,
region: env.LANGFUSE_S3_EVENT_UPLOAD_REGION,
forcePathStyle: env.LANGFUSE_S3_EVENT_UPLOAD_FORCE_PATH_STYLE === "true",
});
}
return s3StorageServiceClient;
};
export async function upsertClickhouse<
T extends Record<string, unknown>,
>(opts: {
table: "scores" | "traces"; // TODO: Modify eventType logic to support more tables going forward
records: T[];
eventBodyMapper: (body: T) => Record<string, unknown>;
}): Promise<void> {
return await instrumentAsync({ name: "clickhouse-upsert" }, async (span) => {
// https://opentelemetry.io/docs/specs/semconv/database/database-spans/
span.setAttribute("ch.query.table", opts.table);
// drop trailing s and pretend it's always a create.
// Only applicable to scores and traces.
const eventType = `${opts.table.slice(0, -1)}-create`;
// If event upload is enabled, we store all rows in S3 to have a backup
if (env.LANGFUSE_S3_EVENT_UPLOAD_ENABLED === "true") {
if (env.LANGFUSE_S3_EVENT_UPLOAD_BUCKET === undefined) {
throw new Error("S3 event store is enabled but no bucket is set");
}
const s3Client = getS3StorageServiceClient(
env.LANGFUSE_S3_EVENT_UPLOAD_BUCKET,
);
await Promise.all(
opts.records.map((record) => {
s3Client.uploadJson(
`${env.LANGFUSE_S3_EVENT_UPLOAD_PREFIX}${record.project_id}/${getClickhouseEntityType(eventType)}/${record.id}/${randomUUID()}.json`,
[
{
id: randomUUID(),
timestamp: new Date().toISOString(),
type: eventType,
body: opts.eventBodyMapper(record),
},
],
);
}),
);
}
const res = await clickhouseClient.insert({
table: opts.table,
values: opts.records.map((record) => ({
...record,
event_ts: convertDateToClickhouseDateTime(new Date()),
})),
format: "JSONEachRow",
});
// same logic as for prisma. we want to see queries in development
if (env.NODE_ENV === "development") {
logger.info(`clickhouse:insert ${res.query_id} ${opts.table}`);
}
span.setAttribute("ch.queryId", res.query_id);
// add summary headers to the span. Helps to tune performance
const summaryHeader = res.response_headers["x-clickhouse-summary"];
if (summaryHeader) {
try {
const summary = Array.isArray(summaryHeader)
? JSON.parse(summaryHeader[0])
: JSON.parse(summaryHeader);
for (const key in summary) {
span.setAttribute(`ch.${key}`, summary[key]);
}
} catch (error) {
logger.debug(
`Failed to parse clickhouse summary header ${summaryHeader}`,
error,
);
}
}
});
}
export async function queryClickhouse<T>(opts: {
query: string;
params?: Record<string, unknown> | undefined;
}): Promise<T[]> {
return await instrumentAsync({ name: "clickhouse-query" }, async (span) => {
// https://opentelemetry.io/docs/specs/semconv/database/database-spans/
span.setAttribute("ch.query.text", opts.query);
const res = await clickhouseClient.query({
query: opts.query,
format: "JSONEachRow",
query_params: opts.params,
});
// same logic as for prisma. we want to see queries in development
if (env.NODE_ENV === "development") {
logger.info(`clickhouse:query ${res.query_id} ${opts.query}`);
}
span.setAttribute("ch.queryId", res.query_id);
// add summary headers to the span. Helps to tune performance
const summaryHeader = res.response_headers["x-clickhouse-summary"];
if (summaryHeader) {
try {
const summary = Array.isArray(summaryHeader)
? JSON.parse(summaryHeader[0])
: JSON.parse(summaryHeader);
for (const key in summary) {
span.setAttribute(`ch.${key}`, summary[key]);
}
} catch (error) {
logger.debug(
`Failed to parse clickhouse summary header ${summaryHeader}`,
error,
);
}
}
return await res.json<T>();
});
}
export async function commandClickhouse<T>(opts: {
query: string;
params?: Record<string, unknown> | undefined;
}): Promise<void> {
return await instrumentAsync({ name: "clickhouse-command" }, async (span) => {
// https://opentelemetry.io/docs/specs/semconv/database/database-spans/
span.setAttribute("ch.query.text", opts.query);
const res = await clickhouseClient.command({
query: opts.query,
query_params: opts.params,
});
// same logic as for prisma. we want to see queries in development
if (env.NODE_ENV === "development") {
logger.info(`clickhouse:query ${res.query_id} ${opts.query}`);
}
span.setAttribute("ch.queryId", res.query_id);
// add summary headers to the span. Helps to tune performance
const summaryHeader = res.response_headers["x-clickhouse-summary"];
if (summaryHeader) {
try {
const summary = Array.isArray(summaryHeader)
? JSON.parse(summaryHeader[0])
: JSON.parse(summaryHeader);
for (const key in summary) {
span.setAttribute(`ch.${key}`, summary[key]);
}
} catch (error) {
logger.debug(
`Failed to parse clickhouse summary header ${summaryHeader}`,
error,
);
}
}
});
}
export function parseClickhouseUTCDateTimeFormat(dateStr: string): Date {
return new Date(`${dateStr.replace(" ", "T")}Z`);
}
@@ -0,0 +1,7 @@
// t.timestamp > observation.start_time - 2 days
export const OBSERVATIONS_TO_TRACE_INTERVAL = "INTERVAL 2 DAY";
// observation.start_time > t.timestamp - 1 hour
export const TRACE_TO_OBSERVATIONS_INTERVAL = "INTERVAL 1 HOUR";
// observation.start_time > s.timestamp - 1 hour
// t.timestamp > s.timestamp - 1 hour
export const SCORE_TO_TRACE_OBSERVATIONS_INTERVAL = "INTERVAL 1 HOUR";
@@ -0,0 +1,556 @@
import { queryClickhouse } from "./clickhouse";
import { createFilterFromFilterState } from "../queries/clickhouse-sql/factory";
import { FilterState } from "../../types";
import {
DateTimeFilter,
FilterList,
} from "../queries/clickhouse-sql/clickhouse-filter";
import { dashboardColumnDefinitions } from "../../tableDefinitions/mapDashboards";
import { convertDateToClickhouseDateTime } from "../clickhouse/client";
import {
SCORE_TO_TRACE_OBSERVATIONS_INTERVAL,
TRACE_TO_OBSERVATIONS_INTERVAL,
} from "./constants";
export type DateTrunc = "year" | "month" | "week" | "day" | "hour" | "minute";
export const getTotalTraces = async (
projectId: string,
filter: FilterState,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
).apply();
const query = `
SELECT
count(id) as count
FROM traces t FINAL
WHERE project_id = {projectId: String}
AND ${chFilter.query}`;
const result = await queryClickhouse<{ count: number }>({
query,
params: {
projectId,
...chFilter.params,
},
});
if (result.length === 0) {
return undefined;
}
return [{ countTraceId: result[0].count }];
};
export const getObservationsCostGroupedByName = async (
projectId: string,
filter: FilterState,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
);
const appliedFilter = chFilter.apply();
const hasTraceFilter = chFilter.find((f) => f.clickhouseTable === "traces");
const query = `
SELECT
provided_model_name as name,
sumMap(cost_details)['total'] as sum_cost_details,
sumMap(usage_details)['total'] as sum_usage_details
FROM observations o FINAL ${hasTraceFilter ? "LEFT JOIN traces t ON o.trace_id = t.id AND o.project_id = t.project_id" : ""}
WHERE project_id = {projectId: String}
AND ${appliedFilter.query}
GROUP BY provided_model_name
ORDER BY sumMap(cost_details)['total'] DESC
`;
const result = await queryClickhouse<{
name: string;
sum_cost_details: number;
sum_usage_details: number;
}>({
query,
params: {
projectId,
...appliedFilter.params,
},
});
return result;
};
export const getScoreAggregate = async (
projectId: string,
filter: FilterState,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
);
const timeFilter = chFilter.find(
(f) =>
f.field === "timestamp" && (f.operator === ">=" || f.operator === ">"),
) as DateTimeFilter | undefined;
const chFilterApplied = chFilter.apply();
const hasTraceFilter = chFilter.find((f) => f.clickhouseTable === "traces");
const query = `
SELECT
s.name,
count(*) as count,
avg(s.value) as avg_value,
s.source,
s.data_type
FROM scores s FINAL
${hasTraceFilter ? "JOIN traces t FINAL ON t.id = s.trace_id AND t.project_id = s.project_id" : ""}
WHERE s.project_id = {projectId: String}
AND ${chFilterApplied.query}
${timeFilter && hasTraceFilter ? `AND t.timestamp >= {tracesTimestamp: DateTime} - ${SCORE_TO_TRACE_OBSERVATIONS_INTERVAL}` : ""}
GROUP BY s.name, s.source, s.data_type
ORDER BY count(*) DESC
`;
const result = await queryClickhouse<{
name: string;
count: string;
avg_value: string;
source: string;
data_type: string;
}>({
query,
params: {
projectId,
...chFilterApplied.params,
...(timeFilter
? { tracesTimestamp: convertDateToClickhouseDateTime(timeFilter.value) }
: {}),
},
});
return result;
};
export const groupTracesByTime = async (
projectId: string,
filter: FilterState,
groupBy: DateTrunc,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
).apply();
const query = `
SELECT
${selectTimeseriesColumn(groupBy, "timestamp", "timestamp")},
count(*) as count
FROM traces t FINAL
WHERE project_id = {projectId: String}
AND ${chFilter.query}
GROUP BY timestamp
${orderByTimeSeries(groupBy, "timestamp")}
`;
const result = await queryClickhouse<{
timestamp: string;
count: string;
}>({
query,
params: {
projectId,
...chFilter.params,
},
});
return result.map((row) => ({
timestamp: new Date(row.timestamp),
countTraceId: Number(row.count),
}));
};
export const getObservationUsageByTime = async (
projectId: string,
filter: FilterState,
groupBy: DateTrunc,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
);
const appliedFilter = chFilter.apply();
const query = `
SELECT
${selectTimeseriesColumn(groupBy, "start_time", "start_time")},
sumMap(usage_details)['total'] as sum_usage_details,
sumMap(cost_details)['total'] as sum_cost_details,
provided_model_name
FROM observations o FINAL
${chFilter.find((f) => f.clickhouseTable === "traces") ? "LEFT JOIN traces t ON o.trace_id = t.id AND o.project_id = t.project_id" : ""}
WHERE project_id = {projectId: String}
AND ${appliedFilter.query}
GROUP BY start_time, provided_model_name
${orderByTimeSeries(groupBy, "start_time")}
`;
const result = await queryClickhouse<{
start_time: string;
sum_usage_details: string;
sum_cost_details: number;
provided_model_name: string;
}>({
query,
params: {
projectId,
...appliedFilter.params,
},
});
return result.map((row) => ({
start_time: new Date(row.start_time),
sum_usage_details: Number(row.sum_usage_details),
sum_cost_details: row.sum_cost_details,
provided_model_name: row.provided_model_name,
}));
};
export const getDistinctModels = async (
projectId: string,
filter: FilterState,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
);
const appliedFilter = chFilter.apply();
const query = `
SELECT
distinct(provided_model_name) as model
FROM observations o FINAL
${chFilter.find((f) => f.clickhouseTable === "traces") ? "LEFT JOIN traces t ON o.trace_id = t.id AND o.project_id = t.project_id" : ""}
WHERE project_id = {projectId: String}
AND ${appliedFilter.query}
`;
const result = await queryClickhouse<{ model: string }>({
query,
params: {
projectId,
...appliedFilter.params,
},
});
return result;
};
export const getScoresAggregateOverTime = async (
projectId: string,
filter: FilterState,
groupBy: DateTrunc,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
);
const appliedFilter = chFilter.apply();
const traceFilter = chFilter.find((f) => f.clickhouseTable === "traces");
const query = `
SELECT
${selectTimeseriesColumn(groupBy, "timestamp", "timestamp")},
name,
data_type,
source,
AVG(value) as avg_value
FROM scores FINAL
${traceFilter ? "JOIN traces t ON scores.trace_id = t.id AND scores.project_id = t.project_id" : ""}
WHERE project_id = {projectId: String}
AND ${appliedFilter.query}
AND data_type IN ('NUMERIC', 'BOOLEAN')
GROUP BY
timestamp,
name,
data_type,
source
${orderByTimeSeries(groupBy, "timestamp")};
`;
const result = await queryClickhouse<{
timestamp: string;
name: string;
data_type: string;
source: string;
avg_value: number;
}>({
query,
params: {
projectId,
...appliedFilter.params,
},
});
return result.map((row) => ({
scoreTimestamp: new Date(row.timestamp),
scoreName: row.name,
scoreDataType: row.data_type,
scoreSource: row.source,
avgValue: Number(row.avg_value),
}));
};
export const getModelUsageByUser = async (
projectId: string,
filter: FilterState,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
);
const appliedFilter = chFilter.apply();
const timeFilter = chFilter.find(
(f) =>
f.clickhouseTable === "observations" &&
f.field === "start_time" &&
(f.operator === ">=" || f.operator === ">"),
) as DateTimeFilter | undefined;
const query = `
SELECT
sumMap(usage_details)['total'] as sum_usage_details,
sumMap(cost_details)['total'] as sum_cost_details,
user_id
FROM observations o FINAL
JOIN traces t FINAL ON o.trace_id = t.id AND o.project_id = t.project_id
WHERE project_id = {projectId: String}
AND t.user_id IS NOT NULL
AND ${appliedFilter.query}
${timeFilter ? `AND t.timestamp >= {tractTimestamp: DateTime} - ${TRACE_TO_OBSERVATIONS_INTERVAL}` : ""}
GROUP BY user_id
ORDER BY sum_cost_details DESC
`;
const result = await queryClickhouse<{
sum_usage_details: string;
sum_cost_details: number;
user_id: string;
}>({
query,
params: {
projectId,
...appliedFilter.params,
...(timeFilter ? { tractTimestamp: timeFilter.value } : {}),
},
});
return result.map((row) => ({
sumUsageDetails: Number(row.sum_usage_details),
sumCostDetails: Number(row.sum_cost_details),
userId: row.user_id,
}));
};
export const getObservationLatencies = async (
projectId: string,
filter: FilterState,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
);
const appliedFilter = chFilter.apply();
const query = `
SELECT
quantilesExactLow(0.5, 0.9, 0.95, 0.99)(date_diff('milliseconds', o.start_time, o.end_time)) as quantiles,
name
FROM observations o FINAL
${chFilter.find((f) => f.clickhouseTable === "traces") ? "LEFT JOIN traces t ON o.trace_id = t.id AND o.project_id = t.project_id" : ""}
WHERE project_id = {projectId: String}
AND ${appliedFilter.query}
GROUP BY name
ORDER BY quantiles[2] DESC
`;
const result = await queryClickhouse<{ quantiles: string[]; name: string }>({
query,
params: { projectId, ...appliedFilter.params },
});
return result.map((row) => ({
p50: Number(row.quantiles[0]) / 1000,
p90: Number(row.quantiles[1]) / 1000,
p95: Number(row.quantiles[2]) / 1000,
p99: Number(row.quantiles[3]) / 1000,
name: row.name,
}));
};
export const getTracesLatencies = async (
projectId: string,
filter: FilterState,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
);
const appliedFilter = chFilter.apply();
const timestampFilter = chFilter.find(
(f) =>
f.clickhouseTable === "traces" &&
f.field === 't."timestamp"' &&
(f.operator === ">=" || f.operator === ">"),
) as DateTimeFilter | undefined;
const query = `
WITH trace_latencies as (
select o.trace_id,
t.name,
o.project_id,
date_diff('milliseconds', min(o.start_time), coalesce(max(o.end_time), max(o.start_time))) as duration
FROM traces t FINAL
JOIN observations o FINAL
ON o.trace_id = t.id AND o.project_id = t.project_id
WHERE project_id = {projectId: String}
AND ${appliedFilter.query}
${timestampFilter ? `AND o.start_time > {dateTimeFilterObservations: DateTime64(3)} - ${TRACE_TO_OBSERVATIONS_INTERVAL}` : ""}
GROUP BY o.project_id, o.trace_id, t.name
)
SELECT
quantilesExactLow(0.5, 0.9, 0.95, 0.99)(duration) as quantiles,
name
FROM trace_latencies
GROUP BY name
ORDER BY quantiles[2] DESC
`;
const result = await queryClickhouse<{ quantiles: string[]; name: string }>({
query,
params: {
projectId,
...appliedFilter.params,
...(timestampFilter
? { dateTimeFilterObservations: timestampFilter.value }
: {}),
},
});
return result.map((row) => ({
p50: Number(row.quantiles[0]) / 1000,
p90: Number(row.quantiles[1]) / 1000,
p95: Number(row.quantiles[2]) / 1000,
p99: Number(row.quantiles[3]) / 1000,
name: row.name,
}));
};
export const getModelLatenciesOverTime = async (
projectId: string,
filter: FilterState,
groupBy: DateTrunc,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(filter, dashboardColumnDefinitions),
);
const appliedFilter = chFilter.apply();
const traceFilter = chFilter.find((f) => f.clickhouseTable === "traces");
const query = `
SELECT
${selectTimeseriesColumn(groupBy, "o.start_time", "start_time_bucket")},
provided_model_name,
quantilesExactLow(0.5, 0.75, 0.9, 0.95, 0.99)(date_diff('milliseconds', o.start_time, o.end_time)) as quantiles
FROM observations o FINAL
${traceFilter ? "JOIN traces t ON o.trace_id = t.id AND o.project_id = t.project_id" : ""}
WHERE project_id = {projectId: String}
AND ${appliedFilter.query}
GROUP BY provided_model_name, start_time_bucket
${orderByTimeSeries(groupBy, "start_time_bucket")};
`;
const result = await queryClickhouse<{
start_time_bucket: string;
provided_model_name: string;
quantiles: string[];
}>({ query, params: { projectId, ...appliedFilter.params } });
return result.map((row) => ({
p50: Number(row.quantiles[0]) / 1000,
p75: Number(row.quantiles[1]) / 1000,
p90: Number(row.quantiles[2]) / 1000,
p95: Number(row.quantiles[3]) / 1000,
p99: Number(row.quantiles[4]) / 1000,
model: row.provided_model_name,
start_time: new Date(row.start_time_bucket),
}));
};
const orderByTimeSeries = (dateTrunc: DateTrunc, col: string) => {
let interval;
switch (dateTrunc) {
case "year":
interval = "toIntervalYear(1)";
break;
case "month":
interval = "toIntervalMonth(1)";
break;
case "week":
interval = "toIntervalWeek(1)";
break;
case "day":
interval = "toIntervalDay(1)";
break;
case "hour":
interval = "toIntervalHour(1)";
break;
case "minute":
interval = "toIntervalMinute(1)";
break;
default:
return undefined;
}
return `ORDER BY ${col} ASC WITH FILL STEP ${interval}`;
};
const selectTimeseriesColumn = (
dateTrunc: DateTrunc,
col: string,
as: String,
) => {
let interval;
switch (dateTrunc) {
case "year":
interval = "toStartOfYear";
break;
case "month":
interval = "toStartOfMonth";
break;
case "week":
interval = "toStartOfWeek";
break;
case "day":
interval = "toStartOfDay";
break;
case "hour":
interval = "toStartOfHour";
break;
case "minute":
interval = "toStartOfMinute";
break;
default:
return undefined;
}
return `${interval}(${col}) as ${as}`;
};
@@ -0,0 +1,338 @@
import z from "zod";
export const clickhouseStringDateSchema = z
.string()
// clickhouse stores UTC like '2024-05-23 18:33:41.602000'
// we need to convert it to '2024-05-23T18:33:41.602000Z'
.transform((str) => str.replace(" ", "T") + "Z")
.pipe(z.string().datetime());
//https://clickhouse.com/docs/en/integrations/javascript#integral-types-int64-int128-int256-uint64-uint128-uint256
// clickhouse returns int64 as string
export const UsageCostSchema = z
.record(z.string(), z.coerce.string().nullable())
.transform((val, ctx) => {
const result: Record<string, number> = {};
for (const key in val) {
if (val[key] !== null && val[key] !== undefined) {
const parsed = Number(val[key]);
if (isNaN(parsed)) {
ctx.addIssue({
code: z.ZodIssueCode.custom,
message: `Key ${key} is not a number`,
});
} else {
result[key] = parsed;
}
}
}
return result;
});
export type UsageCostType = z.infer<typeof UsageCostSchema>;
export const observationRecordBaseSchema = z.object({
id: z.string(),
trace_id: z.string().nullish(),
project_id: z.string(),
type: z.string(),
parent_observation_id: z.string().nullish(),
name: z.string().nullish(),
metadata: z.record(z.string()),
level: z.string().nullish(),
status_message: z.string().nullish(),
version: z.string().nullish(),
input: z.string().nullish(),
output: z.string().nullish(),
provided_model_name: z.string().nullish(),
internal_model_id: z.string().nullish(),
model_parameters: z.string().nullish(),
total_cost: z.number().nullish(),
prompt_id: z.string().nullish(),
prompt_name: z.string().nullish(),
prompt_version: z.number().nullish(),
is_deleted: z.number(),
});
export type ObservationRecordBaseType = z.infer<
typeof observationRecordBaseSchema
>;
export const observationRecordReadSchema = observationRecordBaseSchema.extend({
created_at: clickhouseStringDateSchema,
updated_at: clickhouseStringDateSchema,
start_time: clickhouseStringDateSchema,
end_time: clickhouseStringDateSchema.nullish(),
completion_start_time: clickhouseStringDateSchema.nullish(),
event_ts: clickhouseStringDateSchema,
provided_usage_details: UsageCostSchema,
provided_cost_details: UsageCostSchema,
usage_details: UsageCostSchema,
cost_details: UsageCostSchema,
});
export type ObservationRecordReadType = z.infer<
typeof observationRecordReadSchema
>;
export const observationRecordInsertSchema = observationRecordBaseSchema.extend(
{
created_at: z.number(),
updated_at: z.number(),
start_time: z.number(),
end_time: z.number().nullish(),
completion_start_time: z.number().nullish(),
event_ts: z.number(),
provided_usage_details: UsageCostSchema,
provided_cost_details: UsageCostSchema,
usage_details: UsageCostSchema,
cost_details: UsageCostSchema,
},
);
export type ObservationRecordInsertType = z.infer<
typeof observationRecordInsertSchema
>;
export const traceRecordBaseSchema = z.object({
id: z.string(),
name: z.string().nullish(),
user_id: z.string().nullish(),
metadata: z.record(z.string()),
release: z.string().nullish(),
version: z.string().nullish(),
project_id: z.string(),
public: z.boolean(),
bookmarked: z.boolean(),
tags: z.array(z.string()),
input: z.string().nullish(),
output: z.string().nullish(),
session_id: z.string().nullish(),
is_deleted: z.number(),
});
export type TraceRecordBaseType = z.infer<typeof traceRecordBaseSchema>;
export const traceRecordReadSchema = traceRecordBaseSchema.extend({
timestamp: clickhouseStringDateSchema,
created_at: clickhouseStringDateSchema,
updated_at: clickhouseStringDateSchema,
event_ts: clickhouseStringDateSchema,
});
export type TraceRecordReadType = z.infer<typeof traceRecordReadSchema>;
export const traceRecordInsertSchema = traceRecordBaseSchema.extend({
timestamp: z.number(),
created_at: z.number(),
updated_at: z.number(),
event_ts: z.number(),
});
export type TraceRecordInsertType = z.infer<typeof traceRecordInsertSchema>;
export const scoreRecordBaseSchema = z.object({
id: z.string(),
project_id: z.string(),
trace_id: z.string(),
observation_id: z.string().nullish(),
name: z.string().nullish(),
value: z.union([z.number(), z.string()]).nullish(),
source: z.string(),
comment: z.string().nullish(),
author_user_id: z.string().nullish(),
config_id: z.string().nullish(),
data_type: z.enum(["NUMERIC", "CATEGORICAL", "BOOLEAN"]).nullish(),
string_value: z.string().nullish(),
queue_id: z.string().nullish(),
is_deleted: z.number(),
});
export type ScoreRecordBaseType = z.infer<typeof scoreRecordBaseSchema>;
export const scoreRecordReadSchema = scoreRecordBaseSchema.extend({
created_at: clickhouseStringDateSchema,
updated_at: clickhouseStringDateSchema,
timestamp: clickhouseStringDateSchema,
event_ts: clickhouseStringDateSchema,
});
export type ScoreRecordReadType = z.infer<typeof scoreRecordReadSchema>;
export const scoreRecordInsertSchema = scoreRecordBaseSchema.extend({
created_at: z.number(),
updated_at: z.number(),
timestamp: z.number(),
event_ts: z.number(),
});
export type ScoreRecordInsertType = z.infer<typeof scoreRecordInsertSchema>;
export const convertTraceReadToInsert = (
record: TraceRecordReadType,
): TraceRecordInsertType => {
return {
...record,
created_at: new Date(record.created_at).getTime(),
updated_at: new Date(record.updated_at).getTime(),
timestamp: new Date(record.timestamp).getTime(),
event_ts: new Date(record.event_ts).getTime(),
};
};
export const convertObservationReadToInsert = (
record: ObservationRecordReadType,
): ObservationRecordInsertType => {
const convertDate = (date: string) => new Date(date).getTime();
return {
...record,
created_at: convertDate(record.created_at),
updated_at: convertDate(record.updated_at),
start_time: convertDate(record.start_time),
end_time: record.end_time ? convertDate(record.end_time) : undefined,
completion_start_time: record.completion_start_time
? convertDate(record.completion_start_time)
: undefined,
event_ts: convertDate(record.event_ts),
provided_usage_details: record.provided_usage_details,
provided_cost_details: record.provided_cost_details,
usage_details: record.usage_details,
cost_details: record.cost_details,
};
};
export const convertScoreReadToInsert = (
record: ScoreRecordReadType,
): ScoreRecordInsertType => {
return {
...record,
created_at: new Date(record.created_at).getTime(),
updated_at: new Date(record.updated_at).getTime(),
timestamp: new Date(record.timestamp).getTime(),
event_ts: new Date(record.event_ts).getTime(),
};
};
/**
* Expects a single record from a `select * from traces` query. Must be a raw query to keep original
* column names, not the Prisma mapped names.
*/
export const convertPostgresTraceToInsert = (
trace: Record<string, any>,
): TraceRecordInsertType => {
return {
id: trace.id,
timestamp: trace.timestamp?.getTime(),
name: trace.name,
user_id: trace.user_id,
metadata:
typeof trace.metadata === "string"
? { metadata: trace.metadata }
: Array.isArray(trace.metadata)
? { metadata: trace.metadata }
: trace.metadata,
release: trace.release,
version: trace.version,
project_id: trace.project_id,
public: trace.public,
bookmarked: trace.bookmarked,
tags: trace.tags,
input: trace.input ? JSON.stringify(trace.input) : null,
output: trace.output ? JSON.stringify(trace.output) : null,
session_id: trace.session_id,
created_at: trace.created_at?.getTime(),
updated_at: trace.updated_at?.getTime(),
event_ts: trace.timestamp?.getTime(),
is_deleted: 0,
};
};
/**
* Expects a single record from a
* `select o.*,
* o."modelParameters" as model_parameters,
* p.name as prompt_name,
* p.version as prompt_version
* from observations o
* LEFT JOIN prompts p ON p.id = o.prompt_id`
* query. Must be a raw query to keep original
* column names, not the Prisma mapped names.
*/
export const convertPostgresObservationToInsert = (
observation: Record<string, any>,
): ObservationRecordInsertType => {
return {
id: observation.id,
trace_id: observation.trace_id,
project_id: observation.project_id,
type: observation.type,
parent_observation_id: observation.parent_observation_id,
start_time: observation.start_time?.getTime(),
end_time: observation.end_time?.getTime(),
name: observation.name,
metadata:
typeof observation.metadata === "string"
? { metadata: observation.metadata }
: Array.isArray(observation.metadata)
? { metadata: observation.metadata }
: observation.metadata,
level: observation.level,
status_message: observation.status_message,
version: observation.version,
input: observation.input ? JSON.stringify(observation.input) : null,
output: observation.output ? JSON.stringify(observation.output) : null,
provided_model_name: observation.model,
internal_model_id: observation.internal_model_id,
model_parameters: observation.model_parameters
? JSON.stringify(observation.model_parameters)
: null,
provided_usage_details: {},
usage_details: {
input: observation.prompt_tokens >= 0 ? observation.prompt_tokens : null,
output:
observation.completion_tokens >= 0
? observation.completion_tokens
: null,
total: observation.total_tokens >= 0 ? observation.total_tokens : null,
},
provided_cost_details: {
input: observation.input_cost?.toNumber() ?? null,
output: observation.output_cost?.toNumber() ?? null,
total: observation.total_cost?.toNumber() ?? null,
},
cost_details: {
input: observation.calculated_input_cost?.toNumber() ?? null,
output: observation.calculated_output_cost?.toNumber() ?? null,
total: observation.calculated_total_cost?.toNumber() ?? null,
},
total_cost: observation.calculated_total_cost?.toNumber() ?? null,
completion_start_time: observation.completion_start_time?.getTime(),
prompt_id: observation.prompt_id,
prompt_name: observation.prompt_name,
prompt_version: observation.prompt_version,
created_at: observation.created_at?.getTime(),
updated_at: observation.updated_at?.getTime(),
event_ts: observation.start_time?.getTime(),
is_deleted: 0,
};
};
/**
* Expects a single record from a `select * from scores` query. Must be a raw query to keep original
* column names, not the Prisma mapped names.
*/
export const convertPostgresScoreToInsert = (
score: Record<string, any>,
): ScoreRecordInsertType => {
return {
id: score.id,
timestamp: score.timestamp?.getTime(),
project_id: score.project_id,
trace_id: score.trace_id,
observation_id: score.observation_id,
name: score.name,
value: score.value,
source: score.source,
comment: score.comment,
author_user_id: score.author_user_id,
config_id: score.config_id,
data_type: score.data_type,
string_value: score.string_value,
queue_id: score.queue_id,
created_at: score.created_at?.getTime(),
updated_at: score.updated_at?.getTime(),
event_ts: score.timestamp?.getTime(),
is_deleted: 0,
};
};
@@ -0,0 +1,7 @@
export * from "./scores";
export * from "./traces";
export * from "./observations";
export * from "./types";
export * from "./dashboards";
export * from "./traces_converters";
export * from "./scores_converters";
@@ -0,0 +1,612 @@
import { commandClickhouse, queryClickhouse } from "./clickhouse";
import { Observation, ObservationLevel } from "@prisma/client";
import { logger } from "../logger";
import { InternalServerError, LangfuseNotFoundError } from "../../errors";
import { prisma } from "../../db";
import { ObservationRecordReadType } from "./definitions";
import { FilterState } from "../../types";
import {
DateTimeFilter,
FilterList,
StringFilter,
} from "../queries/clickhouse-sql/clickhouse-filter";
import { FullObservations } from "../queries/createGenerationsQuery";
import { createFilterFromFilterState } from "../queries/clickhouse-sql/factory";
import {
observationsTableTraceUiColumnDefinitions,
observationsTableUiColumnDefinitions,
} from "../../tableDefinitions";
import { TableCount } from "./types";
import { orderByToClickhouseSql } from "../queries/clickhouse-sql/orderby-factory";
import { OrderByState } from "../../interfaces/orderBy";
import { getTracesByIds } from "./traces";
import { convertDateToClickhouseDateTime } from "../clickhouse/client";
import {
convertObservationToView,
convertObservation,
} from "./observations_converters";
import { clickhouseSearchCondition } from "../queries/clickhouse-sql/search";
import { OBSERVATIONS_TO_TRACE_INTERVAL } from "./constants";
export const getObservationsViewForTrace = async (
traceId: string,
projectId: string,
fetchWithInputOutput: boolean = false,
) => {
const query = `
SELECT
id,
trace_id,
project_id,
type,
parent_observation_id,
start_time,
end_time,
name,
metadata,
level,
status_message,
version,
${fetchWithInputOutput ? "input, output," : ""}
provided_model_name,
internal_model_id,
model_parameters,
provided_usage_details,
usage_details,
provided_cost_details,
cost_details,
total_cost,
completion_start_time,
prompt_id,
prompt_name,
prompt_version,
created_at,
updated_at,
event_ts
FROM observations FINAL WHERE trace_id = {traceId: String} AND project_id = {projectId: String}`;
const records = await queryClickhouse<ObservationRecordReadType>({
query,
params: { traceId, projectId },
});
return await Promise.all(
records.map(async (o) => await convertObservationToView(o)),
);
};
export const getObservationById = async (
id: string,
projectId: string,
fetchWithInputOutput: boolean = false,
) => {
const query = `
SELECT
id,
trace_id,
project_id,
type,
parent_observation_id,
start_time,
end_time,
name,
metadata,
level,
status_message,
version,
${fetchWithInputOutput ? "input, output," : ""}
provided_model_name,
internal_model_id,
model_parameters,
provided_usage_details,
usage_details,
provided_cost_details,
cost_details,
total_cost,
completion_start_time,
prompt_id,
prompt_name,
prompt_version,
created_at,
updated_at,
event_ts
FROM observations WHERE id = {id: String} AND project_id = {projectId: String} ORDER BY event_ts desc LIMIT 1 by id, project_id`;
const records = await queryClickhouse<ObservationRecordReadType>({
query,
params: { id, projectId },
});
const mapped = await Promise.all(
records.map(async (r) => await convertObservation(r)),
);
if (mapped.length === 0) {
throw new LangfuseNotFoundError(`Observation with id ${id} not found`);
}
if (mapped.length > 1) {
logger.error(
`Multiple observations found for id ${id} and project ${projectId}`,
);
throw new InternalServerError(
`Multiple observations found for id ${id} and project ${projectId}`,
);
}
return mapped.shift() as Observation;
};
export type ObservationTableQuery = {
projectId: string;
filter: FilterState;
orderBy?: OrderByState;
searchQuery?: string;
limit?: number;
offset?: number;
selectIOAndMetadata?: boolean;
};
export type ObservationTableRecord = {
id: string;
projectId: string;
traceId: string;
observationCount: number;
providedCostDetails: Record<string, number>;
costDetails: Record<string, number>;
usageDetails: Record<string, number>;
providedUsageDetails: Record<string, number>;
latencyMs: number;
level: ObservationLevel;
scoresAvg: Record<string, number>;
scoresValues: Record<string, number>;
};
export type ObservationsTableQueryResult = ObservationRecordReadType & {
latency?: string;
time_to_first_token?: string;
trace_tags?: string[];
trace_name?: string;
trace_user_id?: string;
};
export const getObservationsTableCount = async (opts: ObservationTableQuery) =>
getObservationsTableInternal<TableCount>({
...opts,
select: "count(*) as count",
});
export const getObservationsTable = async (
opts: ObservationTableQuery,
): Promise<FullObservations> => {
const observationRecords = await getObservationsTableInternal<
Omit<
ObservationsTableQueryResult,
"trace_tags" | "trace_name" | "trace_user_id" | "type"
>
>({
...opts,
select: `
o.id as id,
o.name as name,
o."model_parameters" as model_parameters,
o.start_time as "start_time",
o.end_time as "end_time",
o.trace_id as "trace_id",
o.completion_start_time as "completion_start_time",
o.provided_usage_details as "provided_usage_details",
o.usage_details as "usage_details",
o.provided_cost_details as "provided_cost_details",
o.cost_details as "cost_details",
o.level as level,
o.status_message as "status_message",
o.version as version,
o.parent_observation_id as "parent_observation_id",
o.created_at as "created_at",
o.updated_at as "updated_at",
o.provided_model_name as "provided_model_name",
o.total_cost as "total_cost",
internal_model_id as "internal_model_id",
provided_model_name as "provided_model_name",
if(isNull(end_time), NULL, date_diff('milliseconds', start_time, end_time)) as latency,
if(isNull(completion_start_time), NULL, date_diff('milliseconds', start_time, completion_start_time)) as "time_to_first_token"`,
});
const uniqueModels: string[] = Array.from(
new Set(
observationRecords
.map((r) => r.internal_model_id)
.filter((r): r is string => Boolean(r)),
),
);
const [models, traces] = await Promise.all([
uniqueModels.length > 0
? prisma.model.findMany({
where: {
id: {
in: uniqueModels,
},
OR: [{ projectId: opts.projectId }, { projectId: null }],
},
include: {
Price: true,
},
})
: [],
getTracesByIds(
observationRecords
.map((o) => o.trace_id)
.filter((o): o is string => Boolean(o)),
opts.projectId,
),
]);
return await Promise.all(
observationRecords.map(async (o) => {
const model = models.find((p) => p.id === o.internal_model_id);
const trace = traces.find((t) => t.id === o.trace_id);
return {
...(await convertObservationToView(
{ ...o, type: "GENERATION" },
model,
)),
latency: o.latency ? Number(o.latency) / 1000 : null,
timeToFirstToken: o.time_to_first_token
? Number(o.time_to_first_token) / 1000
: null,
traceName: trace?.name ?? null,
traceTags: trace?.tags ?? [],
userId: trace?.userId ?? null,
};
}),
);
};
const getObservationsTableInternal = async <T>(
opts: ObservationTableQuery & { select: string },
): Promise<Array<T>> => {
const { projectId, filter, selectIOAndMetadata, limit, offset, orderBy } =
opts;
const selectString = selectIOAndMetadata
? `
${opts.select},
${selectIOAndMetadata ? `o.input, o.output, o.metadata` : ""}
`
: opts.select;
const scoresFilter = new FilterList([
new StringFilter({
clickhouseTable: "scores",
field: "project_id",
operator: "=",
value: projectId,
}),
]);
const timeFilter = opts.filter.find(
(f) =>
f.column === "Start Time" && (f.operator === ">=" || f.operator === ">"),
);
// query optimisation: joining traces onto observations is expensive. Hence, only join if the UI table contains filters on traces.
const traceTableFilter = opts.filter.filter(
(f) =>
observationsTableTraceUiColumnDefinitions
.map((c) => c.uiTableId)
.includes(f.column) ||
observationsTableTraceUiColumnDefinitions
.map((c) => c.uiTableName)
.includes(f.column),
);
timeFilter
? scoresFilter.push(
new DateTimeFilter({
clickhouseTable: "scores",
field: "timestamp",
operator: ">=",
value: timeFilter.value as Date,
}),
)
: undefined;
const observationsFilter = new FilterList([
new StringFilter({
clickhouseTable: "observations",
field: "project_id",
operator: "=",
value: projectId,
tablePrefix: "o",
}),
]);
observationsFilter.push(
...createFilterFromFilterState(
filter,
observationsTableUiColumnDefinitions,
),
);
const appliedScoresFilter = scoresFilter.apply();
const appliedObservationsFilter = observationsFilter.apply();
const search = clickhouseSearchCondition(opts.searchQuery);
const scoresCte = `WITH scores_avg AS (
SELECT
trace_id,
observation_id,
groupArray(tuple(name, avg_value)) AS "scores_avg"
FROM (
SELECT
trace_id,
observation_id,
name,
avg(value) avg_value,
comment
FROM
scores final
WHERE ${appliedScoresFilter.query}
GROUP BY
trace_id,
observation_id,
name,
comment
ORDER BY
trace_id
) tmp
GROUP BY
trace_id,
observation_id
)`;
if (traceTableFilter.length > 0) {
// joins with traces are very expensive. We need to filter by time as well.
// We assume that a trace has to have been within the last 2 days to be relevant.
const query = `
${scoresCte}
SELECT
${selectString}
FROM observations o FINAL
LEFT JOIN traces t FINAL ON t.id = o.trace_id AND t.project_id = o.project_id
LEFT JOIN scores_avg AS s_avg ON s_avg.trace_id = o.trace_id and s_avg.observation_id = o.id
WHERE ${appliedObservationsFilter.query}
AND o.type = 'GENERATION'
${timeFilter ? `AND t.timestamp > {tracesTimestampFilter: DateTime} - ${OBSERVATIONS_TO_TRACE_INTERVAL}` : ""}
${search.query}
${orderByToClickhouseSql(orderBy ?? null, observationsTableUiColumnDefinitions)}
${limit !== undefined && offset !== undefined ? `LIMIT ${limit} OFFSET ${offset}` : ""};`;
const res = await queryClickhouse<T>({
query,
params: {
...appliedScoresFilter.params,
...appliedObservationsFilter.params,
...(timeFilter
? {
tracesTimestampFilter: convertDateToClickhouseDateTime(
timeFilter.value as Date,
),
}
: {}),
...search.params,
},
});
return res;
} else {
// we query by T, which could also be {count: string}.
const query = `
${scoresCte}
SELECT
${selectString}
FROM observations o FINAL
LEFT JOIN scores_avg AS s_avg ON s_avg.trace_id = o.trace_id and s_avg.observation_id = o.id
WHERE ${appliedObservationsFilter.query}
${orderByToClickhouseSql(orderBy ?? null, observationsTableUiColumnDefinitions)}
${limit !== undefined && offset !== undefined ? `LIMIT ${limit} OFFSET ${offset}` : ""};`;
const res = await queryClickhouse<T>({
query,
params: {
...appliedScoresFilter.params,
...appliedObservationsFilter.params,
},
});
return res;
}
};
export const getObservationsGroupedByModel = async (
projectId: string,
filter: FilterState,
) => {
const observationsFilter = new FilterList([
new StringFilter({
clickhouseTable: "observations",
field: "project_id",
operator: "=",
value: projectId,
tablePrefix: "o",
}),
]);
observationsFilter.push(
...createFilterFromFilterState(
filter,
observationsTableUiColumnDefinitions,
),
);
const appliedObservationsFilter = observationsFilter.apply();
const query = `
SELECT
o.provided_model_name as name
FROM observations o FINAL
WHERE ${appliedObservationsFilter.query}
AND o.type = 'GENERATION'
GROUP BY o.provided_model_name
ORDER BY count() DESC
LIMIT 1000;
`;
const res = await queryClickhouse<{ name: string }>({
query,
params: {
...appliedObservationsFilter.params,
},
});
return res.map((r) => ({ model: r.name }));
};
export const getObservationsGroupedByName = async (
projectId: string,
filter: FilterState,
) => {
const observationsFilter = new FilterList([
new StringFilter({
clickhouseTable: "observations",
field: "project_id",
operator: "=",
value: projectId,
tablePrefix: "o",
}),
]);
observationsFilter.push(
...createFilterFromFilterState(
filter,
observationsTableUiColumnDefinitions,
),
);
const appliedObservationsFilter = observationsFilter.apply();
const query = `
SELECT
o.name as name
FROM observations o FINAL
WHERE ${appliedObservationsFilter.query}
AND o.type = 'GENERATION'
GROUP BY o.name
ORDER BY count() DESC
LIMIT 1000;
`;
const res = await queryClickhouse<{ name: string }>({
query,
params: {
...appliedObservationsFilter.params,
},
});
return res;
};
export const getObservationsGroupedByPromptName = async (
projectId: string,
filter: FilterState,
) => {
const observationsFilter = new FilterList([
new StringFilter({
clickhouseTable: "observations",
field: "project_id",
operator: "=",
value: projectId,
tablePrefix: "o",
}),
]);
observationsFilter.push(
...createFilterFromFilterState(
filter,
observationsTableUiColumnDefinitions,
),
);
const appliedObservationsFilter = observationsFilter.apply();
const query = `
SELECT
o.prompt_id as id
FROM observations o FINAL
WHERE ${appliedObservationsFilter.query}
AND o.type = 'GENERATION'
AND o.prompt_id IS NOT NULL
GROUP BY o.prompt_id
ORDER BY count() DESC
LIMIT 1000;
`;
const res = await queryClickhouse<{ id: string }>({
query,
params: {
...appliedObservationsFilter.params,
},
});
const prompts = res.map((r) => r.id).filter((r): r is string => Boolean(r));
const pgPrompts =
prompts.length > 0
? await prisma.prompt.findMany({
select: {
id: true,
name: true,
},
where: {
id: {
in: prompts,
},
projectId,
},
})
: [];
return pgPrompts.map((p) => ({
promptName: p.name,
}));
};
export const getCostForTraces = async (
projectId: string,
traceIds: string[],
) => {
const query = `
SELECT
sum(o.total_cost) as total_cost
FROM observations o FINAL
WHERE o.project_id = {projectId: String}
AND o.trace_id IN ({traceIds: Array(String)});
`;
const res = await queryClickhouse<{ total_cost: string }>({
query,
params: {
projectId,
traceIds,
},
});
return res.length > 0 ? Number(res[0].total_cost) : undefined;
};
export const deleteObservationsByTraceIds = async (
projectId: string,
traceIds: string[],
) => {
const query = `
DELETE FROM observations
WHERE project_id = {projectId: String}
AND trace_id IN ({traceIds: Array(String)});
`;
await commandClickhouse({
query: query,
params: {
projectId,
traceIds,
},
});
};
@@ -0,0 +1,129 @@
import {
Observation,
ObservationView,
Model,
Price,
ObservationType,
ObservationLevel,
} from "@prisma/client";
import Decimal from "decimal.js";
import { prisma } from "../../db";
import { jsonSchema } from "../../utils/zod";
import { parseClickhouseUTCDateTimeFormat } from "./clickhouse";
import { ObservationRecordReadType } from "./definitions";
export const convertObservation = async (
record: ObservationRecordReadType,
): Promise<Observation> => {
const model = record.internal_model_id
? await prisma.model.findFirst({
where: {
id: record.internal_model_id,
},
include: {
Price: true,
},
})
: undefined;
return convertObservationAndModel(record, model ?? undefined);
};
export const convertObservationToView = async (
record: ObservationRecordReadType,
providedModel?: Model & { Price: Price[] },
): Promise<ObservationView> => {
const model =
providedModel ??
(record.internal_model_id
? await prisma.model.findFirst({
where: {
id: record.internal_model_id,
},
include: {
Price: true,
},
})
: undefined);
return {
...convertObservationAndModel(record, model ?? undefined),
latency: record.end_time
? parseClickhouseUTCDateTimeFormat(record.end_time).getTime() -
parseClickhouseUTCDateTimeFormat(record.start_time).getTime()
: null,
timeToFirstToken: record.completion_start_time
? parseClickhouseUTCDateTimeFormat(record.start_time).getTime() -
parseClickhouseUTCDateTimeFormat(record.completion_start_time).getTime()
: null,
inputPrice:
model?.Price?.find((m) => m.usageType === "input")?.price ?? null,
outputPrice:
model?.Price?.find((m) => m.usageType === "output")?.price ?? null,
totalPrice:
model?.Price?.find((m) => m.usageType === "total")?.price ?? null,
promptName: record.prompt_name ?? null,
promptVersion: record.prompt_version ?? null,
modelId: record.internal_model_id ?? null,
};
};
export const convertObservationAndModel = (
record: ObservationRecordReadType,
model?: Model & { Price: Price[] },
): Observation => {
return {
id: record.id,
traceId: record.trace_id ?? null,
projectId: record.project_id,
type: record.type as ObservationType,
parentObservationId: record.parent_observation_id ?? null,
startTime: parseClickhouseUTCDateTimeFormat(record.start_time),
endTime: record.end_time
? parseClickhouseUTCDateTimeFormat(record.end_time)
: null,
name: record.name ?? null,
metadata: record.metadata,
level: record.level as ObservationLevel,
statusMessage: record.status_message ?? null,
version: record.version ?? null,
input: jsonSchema.nullish().parse(record.input) ?? null,
output: jsonSchema.nullish().parse(record.output) ?? null,
modelParameters: record.model_parameters
? JSON.parse(record.model_parameters)
: null,
completionStartTime: record.completion_start_time
? parseClickhouseUTCDateTimeFormat(record.completion_start_time)
: null,
promptId: record.prompt_id ?? null,
createdAt: parseClickhouseUTCDateTimeFormat(record.created_at),
updatedAt: parseClickhouseUTCDateTimeFormat(record.updated_at),
promptTokens: record.usage_details?.input
? Number(record.usage_details?.input)
: 0,
completionTokens: record.usage_details?.output
? Number(record.usage_details?.output)
: 0,
totalTokens: record.usage_details?.total
? Number(record.usage_details?.total)
: 0,
calculatedInputCost: record.cost_details?.input
? new Decimal(record.cost_details.input)
: null,
calculatedOutputCost: record.cost_details?.output
? new Decimal(record.cost_details.output)
: null,
calculatedTotalCost: record.cost_details?.total
? new Decimal(record.cost_details.total)
: null,
inputCost: record.cost_details?.input
? new Decimal(record.cost_details?.input)
: null,
outputCost: record.cost_details?.output
? new Decimal(record.cost_details?.output)
: null,
totalCost: record.total_cost ? new Decimal(record.total_cost) : null,
model: record.provided_model_name ?? null,
internalModelId: record.internal_model_id ?? null,
internalModel: model?.modelName ?? null, // to be removed
unit: "TOKENS", // to be removed.
};
};
@@ -0,0 +1,473 @@
import { Score, ScoreDataType, ScoreSource } from "@prisma/client";
import {
commandClickhouse,
parseClickhouseUTCDateTimeFormat,
queryClickhouse,
upsertClickhouse,
} from "./clickhouse";
import { FilterList } from "../queries/clickhouse-sql/clickhouse-filter";
import { FilterState } from "../../types";
import {
createFilterFromFilterState,
getProjectIdDefaultFilter,
} from "../queries/clickhouse-sql/factory";
import { OrderByState } from "../../interfaces/orderBy";
import { scoresTableUiColumnDefinitions } from "../../tableDefinitions";
import { orderByToClickhouseSql } from "../queries/clickhouse-sql/orderby-factory";
import { convertToScore } from "./scores_converters";
export type FetchScoresReturnType = {
id: string;
timestamp: string;
project_id: string;
trace_id: string;
observation_id: string | null;
name: string;
value: number;
source: string;
comment: string | null;
author_user_id: string | null;
config_id: string | null;
data_type: string;
string_value: string | null;
queue_id: string | null;
created_at: string;
updated_at: string;
event_ts: string;
is_deleted: number;
projectId: string;
};
export const searchExistingAnnotationScore = async (
projectId: string,
traceId: string,
observationId: string | null,
name: string | undefined,
configId: string | undefined,
) => {
if (!name && !configId) {
throw new Error("Either name or configId (or both) must be provided.");
}
const query = `
SELECT *
FROM scores s FINAL
WHERE s.project_id = {projectId: String}
AND s.source = 'ANNOTATION'
AND s.trace_id = {traceId: String}
${observationId ? `AND s.observation_id = {observationId: String}` : "AND isNull(s.observation_id)"}
AND (
FALSE
${name ? `OR s.name = {name: String}` : ""}
${configId ? `OR s.config_id = {configId: String}` : ""}
)
ORDER BY s.event_ts DESC
LIMIT 1
`;
const rows = await queryClickhouse<FetchScoresReturnType>({
query,
params: {
projectId,
name,
configId,
traceId,
observationId,
},
});
return rows.map(convertToScore).shift();
};
export const getScoreById = async (
projectId: string,
scoreId: string,
source: ScoreSource,
) => {
const query = `
SELECT *
FROM scores s FINAL
WHERE s.project_id = {projectId: String}
AND s.id = {scoreId: String}
AND s.source = {source: String}
ORDER BY s.event_ts DESC
LIMIT 1
`;
const rows = await queryClickhouse<FetchScoresReturnType>({
query,
params: {
projectId,
scoreId,
source,
},
});
return rows.map(convertToScore).shift();
};
/**
* Accepts a score in a Clickhouse-ready format.
* id, project_id, name, and timestamp must always be provided.
*/
export const upsertScore = async (score: Partial<FetchScoresReturnType>) => {
if (!["id", "project_id", "name", "timestamp"].every((key) => key in score)) {
throw new Error("Identifier fields must be provided to upsert Score.");
}
await upsertClickhouse({
table: "scores",
records: [score as FetchScoresReturnType],
eventBodyMapper: convertToScore,
});
};
export const getScoresForTraces = async (
projectId: string,
traceIds: string[],
limit?: number,
offset?: number,
) => {
const query = `
select
*
from scores s final
WHERE s.project_id = {projectId: String}
AND s.trace_id IN ({traceIds: Array(String)})
${limit && offset ? `limit {limit: Int32} offset {offset: Int32}` : ""}
`;
const rows = await queryClickhouse<FetchScoresReturnType>({
query: query,
params: {
projectId: projectId,
traceIds: traceIds,
limit: limit,
offset: offset,
},
});
return rows.map(convertToScore);
};
export const getScoresForObservations = async (
projectId: string,
observationIds: string[],
limit?: number,
offset?: number,
) => {
const query = `
select
*
from scores s final
WHERE s.project_id = {projectId: String}
AND s.observation_id IN ({observationIds: Array(String)})
${limit !== undefined && offset !== undefined ? `limit {limit: Int32} offset {offset: Int32}` : ""}
`;
const rows = await queryClickhouse<FetchScoresReturnType>({
query: query,
params: {
projectId: projectId,
observationIds: observationIds,
limit: limit,
offset: offset,
},
});
return rows.map(convertToScore);
};
export const getScoresGroupedByNameSourceType = async (projectId: string) => {
const query = `
select
name,
source,
data_type
from scores s final
WHERE s.project_id = {projectId: String}
GROUP BY name, source, data_type
ORDER BY count() desc
LIMIT 1000;
`;
const rows = await queryClickhouse<{
name: string;
source: string;
data_type: string;
}>({
query: query,
params: {
projectId: projectId,
},
});
return rows.map((row) => ({
name: row.name,
source: row.source as ScoreSource,
dataType: row.data_type as ScoreDataType,
}));
};
export const getScoresGroupedByName = async (
projectId: string,
timestampFilter?: FilterState,
) => {
const chFilter = timestampFilter
? createFilterFromFilterState(timestampFilter, [
{
uiTableName: "Timestamp",
uiTableId: "timestamp",
clickhouseTableName: "scores",
clickhouseSelect: "timestamp",
},
])
: undefined;
const timestampFilterRes = chFilter
? new FilterList(chFilter).apply()
: undefined;
const query = `
select
name as name
from scores s final
WHERE s.project_id = {projectId: String}
AND has(['NUMERIC', 'BOOLEAN'], s.data_type)
${timestampFilterRes?.query ? `AND ${timestampFilterRes.query}` : ""}
GROUP BY name
ORDER BY count() desc
LIMIT 1000;
`;
const rows = await queryClickhouse<{
name: string;
}>({
query: query,
params: {
projectId: projectId,
...(timestampFilterRes ? timestampFilterRes.params : {}),
},
});
return rows;
};
export const getScoresUiCount = async (props: {
projectId: string;
filter: FilterState;
orderBy: OrderByState;
limit?: number;
offset?: number;
}) => {
const rows = await getScoresUiGeneric<{ count: string }>({
select: `
count(*) as count
`,
...props,
});
return Number(rows[0].count);
};
export type ScoreUiTableRow = Score & {
traceName: string | null;
traceUserId: string | null;
traceTags: Array<string> | null;
};
export const getScoresUiTable = async (props: {
projectId: string;
filter: FilterState;
orderBy: OrderByState;
limit?: number;
offset?: number;
}): Promise<ScoreUiTableRow[]> => {
const rows = await getScoresUiGeneric<{
id: string;
project_id: string;
name: string;
value: number;
string_value: string | null;
timestamp: string;
source: string;
data_type: string;
comment: string | null;
trace_id: string;
observation_id: string | null;
author_user_id: string | null;
user_id: string | null;
trace_name: string | null;
trace_tags: Array<string> | null;
job_configuration_id: string | null;
author_user_image: string | null;
author_user_name: string | null;
config_id: string | null;
queue_id: string | null;
created_at: string;
updated_at: string;
}>({
select: `
s.id,
s.project_id,
s.name,
s.value,
s.string_value,
s.timestamp,
s.source,
s.data_type,
s.comment,
s.trace_id,
s.observation_id,
s.author_user_id,
t.user_id,
t.name,
t.tags,
s.created_at,
s.updated_at,
s.source,
s.config_id,
s.queue_id,
t.user_id,
t.name as trace_name,
t.tags as trace_tags
`,
...props,
});
return rows.map((row) => ({
projectId: row.project_id,
authorUserId: row.author_user_id,
traceId: row.trace_id,
observationId: row.observation_id,
traceUserId: row.user_id,
traceName: row.trace_name,
traceTags: row.trace_tags,
configId: row.config_id,
queueId: row.queue_id,
createdAt: parseClickhouseUTCDateTimeFormat(row.created_at),
updatedAt: parseClickhouseUTCDateTimeFormat(row.updated_at),
stringValue: row.string_value,
comment: row.comment,
dataType: row.data_type as ScoreDataType,
source: row.source as ScoreSource,
name: row.name,
value: row.value,
timestamp: parseClickhouseUTCDateTimeFormat(row.timestamp),
id: row.id,
}));
};
export const getScoresUiGeneric = async <T>(props: {
select: string;
projectId: string;
filter: FilterState;
orderBy: OrderByState;
limit?: number;
offset?: number;
}): Promise<T[]> => {
const { select, projectId, filter, orderBy, limit, offset } = props;
const { tracesFilter, scoresFilter, observationsFilter } =
getProjectIdDefaultFilter(projectId, { tracesPrefix: "t" });
scoresFilter.push(
...createFilterFromFilterState(filter, scoresTableUiColumnDefinitions),
);
const scoresFilterRes = scoresFilter.apply();
const query = `
SELECT
${select}
FROM scores s final
LEFT JOIN traces t ON s.trace_id = t.id AND t.project_id = s.project_id
WHERE s.project_id = {projectId: String}
${scoresFilterRes?.query ? `AND ${scoresFilterRes.query}` : ""}
${orderByToClickhouseSql(orderBy ?? null, scoresTableUiColumnDefinitions)}
${limit !== undefined && offset !== undefined ? `limit {limit: Int32} offset {offset: Int32}` : ""}
`;
const rows = await queryClickhouse<T>({
query: query,
params: {
projectId: projectId,
...(scoresFilterRes ? scoresFilterRes.params : {}),
limit: limit,
offset: offset,
},
});
return rows;
};
export const getScoreNames = async (
projectId: string,
timestampFilter: FilterState,
) => {
const chFilter = new FilterList(
createFilterFromFilterState(
timestampFilter,
scoresTableUiColumnDefinitions,
),
);
const timestampFilterRes = chFilter.apply();
const query = `
select
name,
count(*) as count
from scores s final
WHERE s.project_id = {projectId: String}
${timestampFilterRes?.query ? `AND ${timestampFilterRes.query}` : ""}
GROUP BY name
ORDER BY count() desc
LIMIT 1000;
`;
const rows = await queryClickhouse<{
name: string;
count: string;
}>({
query: query,
params: {
projectId: projectId,
...(timestampFilterRes ? timestampFilterRes.params : {}),
},
});
return rows.map((row) => ({
name: row.name,
count: Number(row.count),
}));
};
export const deleteScore = async (projectId: string, scoreId: string) => {
const query = `
DELETE FROM scores
WHERE project_id = {projectId: String}
AND id = {scoreId: String};
`;
await commandClickhouse({
query: query,
params: {
projectId,
scoreId,
},
});
};
export const deleteScoresByTraceIds = async (
projectId: string,
traceIds: string[],
) => {
const query = `
DELETE FROM scores
WHERE project_id = {projectId: String}
AND trace_id IN ({traceIds: Array(String)});
`;
await commandClickhouse({
query: query,
params: {
projectId,
traceIds,
},
});
};
@@ -0,0 +1,23 @@
import { ScoreSource, ScoreDataType } from "@prisma/client";
import { FetchScoresReturnType } from "./scores";
export const convertToScore = (row: FetchScoresReturnType) => {
return {
id: row.id,
timestamp: new Date(row.timestamp),
projectId: row.project_id,
traceId: row.trace_id,
observationId: row.observation_id,
name: row.name,
value: row.value,
source: row.source as ScoreSource,
comment: row.comment,
authorUserId: row.author_user_id,
configId: row.config_id,
dataType: row.data_type as ScoreDataType,
stringValue: row.string_value,
queueId: row.queue_id,
createdAt: new Date(row.created_at),
updatedAt: new Date(row.updated_at),
};
};
@@ -0,0 +1,738 @@
import {
commandClickhouse,
parseClickhouseUTCDateTimeFormat,
queryClickhouse,
upsertClickhouse,
} from "./clickhouse";
import {
createFilterFromFilterState,
getProjectIdDefaultFilter,
} from "../queries/clickhouse-sql/factory";
import { ObservationLevel, Trace } from "@prisma/client";
import { FilterState } from "../../types";
import { logger } from "../logger";
import {
DateTimeFilter,
FilterList,
StringFilter,
StringOptionsFilter,
} from "../queries/clickhouse-sql/clickhouse-filter";
import { TraceRecordReadType } from "./definitions";
import { tracesTableUiColumnDefinitions } from "../../tableDefinitions/mapTracesTable";
import { OrderByState } from "../../interfaces/orderBy";
import { orderByToClickhouseSql } from "../queries/clickhouse-sql/orderby-factory";
import { UiColumnMapping } from "../../tableDefinitions";
import { sessionCols } from "../../tableDefinitions/mapSessionTable";
import { convertDateToClickhouseDateTime } from "../clickhouse/client";
import { convertClickhouseToDomain } from "./traces_converters";
import { clickhouseSearchCondition } from "../queries/clickhouse-sql/search";
import { TRACE_TO_OBSERVATIONS_INTERVAL } from "./constants";
export type TracesTableReturnType = Pick<
TraceRecordReadType,
| "project_id"
| "id"
| "name"
| "timestamp"
| "bookmarked"
| "release"
| "version"
| "user_id"
| "session_id"
| "tags"
| "metadata"
| "public"
> & {
level: ObservationLevel;
observation_count: number | null;
latency_milliseconds: string | null;
usage_details: Record<string, number>;
cost_details: Record<string, number>;
scores_avg: Array<{ name: string; avg_value: number }>;
};
export const getTracesTableCount = async (props: {
projectId: string;
filter: FilterState;
searchQuery?: string;
orderBy?: OrderByState;
limit?: number;
offset?: number;
}) => {
const countRows = await getTracesTableGeneric<{ count: string }>({
select: "count(*) as count",
...props,
});
const converted = countRows.map((row) => ({
count: Number(row.count),
}));
return converted.length > 0 ? converted[0].count : 0;
};
export const getTracesTable = async (
projectId: string,
filter: FilterState,
searchQuery?: string,
orderBy?: OrderByState,
limit?: number,
offset?: number,
) => {
const rows = await getTracesTableGeneric<TracesTableReturnType>({
select: `
t.id,
t.project_id,
t.timestamp,
t.tags,
t.bookmarked,
t.name,
t.release,
t.version,
t.user_id,
t.session_id,
os.latency_milliseconds,
os.cost_details as cost_details,
os.usage_details as usage_details,
os.level as level,
os.observation_count as observation_count,
s.scores_avg as scores_avg,
t.metadata,
t.public`,
projectId,
filter,
searchQuery,
orderBy,
limit,
offset,
});
return rows;
};
type FetchTracesTableProps = {
select: string;
projectId: string;
filter: FilterState;
searchQuery?: string;
orderBy?: OrderByState;
limit?: number;
offset?: number;
};
const getTracesTableGeneric = async <T>(props: FetchTracesTableProps) => {
const { select, projectId, filter, orderBy, limit, offset, searchQuery } =
props;
const { tracesFilter, scoresFilter, observationsFilter } =
getProjectIdDefaultFilter(projectId, { tracesPrefix: "t" });
tracesFilter.push(
...createFilterFromFilterState(filter, tracesTableUiColumnDefinitions),
);
const traceIdFilter = tracesFilter.find(
(f) => f.clickhouseTable === "traces" && f.field === "id",
) as StringFilter | StringOptionsFilter | undefined;
traceIdFilter
? scoresFilter.push(
new StringOptionsFilter({
clickhouseTable: "scores",
field: "trace_id",
operator: "any of",
values:
traceIdFilter instanceof StringFilter
? [traceIdFilter.value]
: traceIdFilter.values,
}),
)
: null;
traceIdFilter
? observationsFilter.push(
new StringOptionsFilter({
clickhouseTable: "observations",
field: "trace_id",
operator: "any of",
values:
traceIdFilter instanceof StringFilter
? [traceIdFilter.value]
: traceIdFilter.values,
}),
)
: null;
// for query optimisation, we have to add the timeseries filter to observations + scores as well
// stats show, that 98% of all observations have their start_time larger than trace.timestamp - 5 min
const timeStampFilter = tracesFilter.find(
(f) =>
f.field === "timestamp" && (f.operator === ">=" || f.operator === ">"),
) as DateTimeFilter | undefined;
timeStampFilter
? scoresFilter.push(
new DateTimeFilter({
clickhouseTable: "scores",
field: "timestamp",
operator: ">=",
value: timeStampFilter.value,
}),
)
: null;
timeStampFilter
? observationsFilter.push(
new DateTimeFilter({
clickhouseTable: "observations",
field: "start_time",
operator: ">=",
value: timeStampFilter.value,
}),
)
: null;
const tracesFilterRes = tracesFilter.apply();
const scoresFilterRes = scoresFilter.apply();
const observationFilterRes = observationsFilter.apply();
const search = clickhouseSearchCondition(searchQuery);
const query = `
WITH observations_stats AS (
SELECT
COUNT(*) AS observation_count,
sumMap(usage_details) as usage_details,
SUM(total_cost) AS total_cost,
date_diff('milliseconds', least(min(start_time), min(end_time)), greatest(max(start_time), max(end_time))) as latency_milliseconds,
multiIf(
arrayExists(x -> x = 'ERROR', groupArray(level)), 'ERROR',
arrayExists(x -> x = 'WARNING', groupArray(level)), 'WARNING',
arrayExists(x -> x = 'DEFAULT', groupArray(level)), 'DEFAULT',
'DEBUG'
) AS level,
sumMap(cost_details) as cost_details,
trace_id,
project_id
FROM
observations final
WHERE ${observationFilterRes.query}
group by trace_id, project_id
),
scores_avg AS (SELECT project_id,
trace_id,
groupArray(tuple(name, avg_value)) AS "scores_avg"
FROM (
SELECT project_id,
trace_id,
name,
avg(value) avg_value
FROM scores final
WHERE ${scoresFilterRes.query}
GROUP BY project_id,
trace_id,
name
) tmp
GROUP BY project_id,
trace_id)
select
${select}
from traces t final
left join observations_stats os on os.project_id = t.project_id and os.trace_id = t.id
left join scores_avg s on s.project_id = t.project_id and s.trace_id = t.id
WHERE ${tracesFilterRes.query}
${search.query}
${orderByToClickhouseSql(orderBy ?? null, tracesTableUiColumnDefinitions)}
${limit !== undefined && offset !== undefined ? `LIMIT {limit: Int32} OFFSET {offset: Int32}` : ""}
`;
const res = await queryClickhouse<T>({
query: query,
params: {
limit: limit,
offset: offset,
...tracesFilterRes.params,
...observationFilterRes.params,
...scoresFilterRes.params,
...search.params,
},
});
return res;
};
/**
* Accepts a trace in a Clickhouse-ready format.
* id, project_id, and timestamp must always be provided.
*/
export const upsertTrace = async (trace: Partial<TraceRecordReadType>) => {
if (!["id", "project_id", "timestamp"].every((key) => key in trace)) {
throw new Error("Identifier fields must be provided to upsert Trace.");
}
await upsertClickhouse({
table: "traces",
records: [trace as TraceRecordReadType],
eventBodyMapper: convertClickhouseToDomain,
});
};
export const getTraceById = async (
traceId: string,
projectId: string,
timestamp?: Date,
): Promise<Trace | undefined> => {
try {
return getTraceByIdOrThrow(traceId, projectId, timestamp);
} catch (e) {
return undefined;
}
};
export const getTracesByIds = async (
traceIds: string[],
projectId: string,
timestamp?: Date,
) => {
const query = `
SELECT *
FROM traces
WHERE id IN ({traceIds: Array(String)})
AND project_id = {projectId: String}
${timestamp ? `AND timestamp >= {timestamp: DateTime}` : ""}
ORDER BY event_ts DESC LIMIT 1 by id, project_id;`;
const records = await queryClickhouse<TraceRecordReadType>({
query,
params: {
traceIds,
projectId,
timestamp: timestamp ? convertDateToClickhouseDateTime(timestamp) : null,
},
});
return records.map(convertClickhouseToDomain);
};
export const getTraceByIdOrThrow = async (
traceId: string,
projectId: string,
timestamp?: Date,
) => {
const query = `SELECT *
FROM traces
WHERE id = {traceId: String}
AND project_id = {projectId: String}
${timestamp ? `AND timestamp = {timestamp: DateTime64(3)}` : ""}
ORDER BY event_ts DESC LIMIT 1 by id, project_id`;
const records = await queryClickhouse<TraceRecordReadType>({
query,
params: {
traceId,
projectId,
timestamp: timestamp ? timestamp.getTime() : null,
},
});
const res = records.map(convertClickhouseToDomain);
if (res.length === 0) {
const errorMessage = `Trace not found for traceId: ${traceId}, projectId: ${projectId}`;
logger.error(errorMessage);
throw new Error(errorMessage);
}
return res[0] as Trace;
};
export const getTracesGroupedByName = async (
projectId: string,
tableDefinitions: UiColumnMapping[] = tracesTableUiColumnDefinitions,
timestampFilter?: FilterState,
) => {
const chFilter = timestampFilter
? createFilterFromFilterState(timestampFilter, tableDefinitions)
: undefined;
const timestampFilterRes = chFilter
? new FilterList(chFilter).apply()
: undefined;
const query = `
select
name as name,
count(*) as count
from traces t final
WHERE t.project_id = {projectId: String}
AND t.name IS NOT NULL
${timestampFilterRes?.query ? `AND ${timestampFilterRes.query}` : ""}
GROUP BY name
ORDER BY count(*) desc
LIMIT 1000;
`;
const rows = await queryClickhouse<{
name: string;
count: string;
}>({
query: query,
params: {
projectId: projectId,
...(timestampFilterRes ? timestampFilterRes.params : {}),
},
});
return rows;
};
export const getTracesGroupedByUsers = async (
projectId: string,
filter: FilterState,
columns?: UiColumnMapping[],
) => {
const chFilter = createFilterFromFilterState(
filter,
columns ?? tracesTableUiColumnDefinitions,
);
const filterRes = new FilterList(chFilter).apply();
const query = `
select
user_id as user,
count(*) as count
from traces t final
WHERE t.project_id = {projectId: String}
AND t.user_id IS NOT NULL
${filterRes?.query ? `AND ${filterRes.query}` : ""}
GROUP BY user
ORDER BY count desc
LIMIT 1000;
`;
const rows = await queryClickhouse<{
user: string;
count: string;
}>({
query: query,
params: {
projectId: projectId,
...(filterRes ? filterRes.params : {}),
},
});
return rows;
};
export type GroupedTracesQueryProp = {
projectId: string;
filter: FilterState;
sessionIdNullFilter?: boolean;
columns?: UiColumnMapping[];
};
export const getTracesGroupedByTags = async (props: GroupedTracesQueryProp) => {
const { projectId, filter, sessionIdNullFilter, columns } = props;
const chFilter = createFilterFromFilterState(
filter,
columns ?? tracesTableUiColumnDefinitions,
);
const filterRes = new FilterList(chFilter).apply();
const query = `
select
distinct(arrayJoin(tags)) as value
from traces t final
WHERE t.project_id = {projectId: String}
${sessionIdNullFilter ? "AND t.session_id IS NOT NULL" : ""}
${filterRes?.query ? `AND ${filterRes.query}` : ""}
LIMIT 1000;
`;
const rows = await queryClickhouse<{
value: string;
}>({
query: query,
params: {
projectId: projectId,
...(filterRes ? filterRes.params : {}),
},
});
return rows;
};
export const getTracesGroupedByUserIds = async (
props: GroupedTracesQueryProp,
) => {
const {
projectId,
filter,
sessionIdNullFilter: sessionIdNotNullFilter,
columns,
} = props;
const chFilter = createFilterFromFilterState(
filter,
columns ?? tracesTableUiColumnDefinitions,
);
const appliedFilter = new FilterList(chFilter).apply();
const query = `
select distinct user_id as user_id
from traces t final
WHERE t.project_id = {projectId: String}
${sessionIdNotNullFilter ? "AND t.session_id IS NOT NULL" : ""}
${appliedFilter?.query ? `AND ${appliedFilter.query}` : ""}
LIMIT 1000;
`;
const rows = await queryClickhouse<{
user_id: string;
}>({
query: query,
params: {
projectId: projectId,
...(appliedFilter ? appliedFilter.params : {}),
},
});
return rows;
};
export type SessionDataReturnType = {
session_id: string;
max_timestamp: string;
min_timestamp: string;
trace_ids: string[];
user_ids: string[];
trace_count: number;
trace_tags: string[];
total_observations: number;
duration: number;
session_usage_details: Record<string, number>;
session_cost_details: Record<string, number>;
session_input_cost: string;
session_output_cost: string;
session_total_cost: string;
session_input_usage: string;
session_output_usage: string;
session_total_usage: string;
};
export const getSessionsTableCount = async (props: {
projectId: string;
filter: FilterState;
orderBy?: OrderByState;
limit?: number;
offset?: number;
}) => {
const rows = await getSessionsTableGeneric<{ count: string }>({
select: `
count(session_id) as count
`,
projectId: props.projectId,
filter: props.filter,
orderBy: props.orderBy,
limit: props.limit,
offset: props.offset,
});
return rows.length > 0 ? Number(rows[0].count) : 0;
};
export const getSessionsTable = async (props: {
projectId: string;
filter: FilterState;
orderBy?: OrderByState;
limit?: number;
offset?: number;
}) => {
const rows = await getSessionsTableGeneric<SessionDataReturnType>({
select: `
session_id,
max_timestamp,
min_timestamp,
trace_ids,
user_ids,
trace_count,
trace_tags,
total_observations,
duration,
session_usage_details,
session_cost_details,
session_input_cost,
session_output_cost,
session_total_cost,
session_input_usage,
session_output_usage,
session_total_usage
`,
projectId: props.projectId,
filter: props.filter,
orderBy: props.orderBy,
limit: props.limit,
offset: props.offset,
});
return rows;
};
const getSessionsTableGeneric = async <T>(props: FetchTracesTableProps) => {
const { select, projectId, filter, orderBy, limit, offset } = props;
const { tracesFilter, scoresFilter, observationsFilter } =
getProjectIdDefaultFilter(projectId, { tracesPrefix: "s" });
tracesFilter.push(...createFilterFromFilterState(filter, sessionCols));
const tracesFilterRes = tracesFilter.apply();
const scoresAvgFilterRes = scoresFilter.apply();
const observationsStatsRes = observationsFilter.apply();
const traceTimestampFilter: DateTimeFilter | undefined = tracesFilter.find(
(f) =>
f.field === "min_timestamp" &&
(f.operator === ">=" || f.operator === ">"),
) as DateTimeFilter | undefined;
const singleTraceFilter = traceTimestampFilter
? new FilterList([
new DateTimeFilter({
clickhouseTable: "traces",
field: "timestamp",
operator: traceTimestampFilter.operator,
value: traceTimestampFilter.value,
}),
]).apply()
: undefined;
const query = `
WITH observations_agg AS (
SELECT o.trace_id,
count(*) as obs_count,
min(o.start_time) as min_start_time,
max(o.end_time) as max_end_time,
sumMap(usage_details) as sum_usage_details,
sumMap(cost_details) as sum_cost_details,
anyLast(project_id) as project_id
FROM observations o FINAL
WHERE o.project_id = {projectId: String}
${traceTimestampFilter ? `AND o.start_time >= {observationsStartTime: DateTime} - ${TRACE_TO_OBSERVATIONS_INTERVAL}` : ""}
GROUP BY o.trace_id
),
session_data AS (
SELECT
t.session_id,
anyLast(t.project_id) as project_id,
max(t.timestamp) as max_timestamp,
min(t.timestamp) as min_timestamp,
groupArray(t.id) AS trace_ids,
groupUniqArray(t.user_id) AS user_ids,
count(*) as trace_count,
groupUniqArrayArray(t.tags) as trace_tags,
-- Aggregate observations data at session level
sum(o.obs_count) as total_observations,
date_diff('milliseconds', min(min_start_time), max(max_end_time)) as duration,
sumMap(o.sum_usage_details) as session_usage_details,
sumMap(o.sum_cost_details) as session_cost_details,
sumMap(o.sum_cost_details)['input'] as session_input_cost,
sumMap(o.sum_cost_details)['output'] as session_output_cost,
sumMap(o.sum_cost_details)['total'] as session_total_cost,
sumMap(o.sum_usage_details)['input'] as session_input_usage,
sumMap(o.sum_usage_details)['output'] as session_output_usage,
sumMap(o.sum_usage_details)['total'] as session_total_usage
FROM traces t FINAL
LEFT JOIN observations_agg o ON t.id = o.trace_id AND t.project_id = o.project_id
WHERE t.session_id IS NOT NULL
AND t.project_id = {projectId: String}
${singleTraceFilter?.query ? ` AND ${singleTraceFilter.query}` : ""}
GROUP BY t.session_id
)
SELECT ${select}
FROM session_data s
WHERE ${tracesFilterRes.query ? tracesFilterRes.query : ""}
${orderByToClickhouseSql(orderBy ?? null, sessionCols)}
${limit !== undefined && offset !== undefined ? `LIMIT {limit: Int32} OFFSET {offset: Int32}` : ""}
`;
const obsStartTimeValue = traceTimestampFilter
? convertDateToClickhouseDateTime(traceTimestampFilter.value)
: null;
const res = await queryClickhouse<T>({
query: query,
params: {
projectId,
limit: limit,
offset: offset,
...tracesFilterRes.params,
...observationsStatsRes.params,
...scoresAvgFilterRes.params,
...singleTraceFilter?.params,
...(obsStartTimeValue
? { observationsStartTime: obsStartTimeValue }
: {}),
},
});
return res;
};
export const getTracesForSession = async (
projectId: string,
sessionId: string,
) => {
const query = `
SELECT
id,
user_id,
name,
timestamp,
project_id
FROM traces
WHERE (project_id = {projectId: String}) AND (session_id = {sessionId: String})
ORDER BY timestamp ASC
LIMIT 1 BY
id,
project_id;
`;
const rows = await queryClickhouse<{
id: string;
user_id: string;
name: string;
timestamp: string;
}>({
query: query,
params: {
projectId,
sessionId,
},
});
return rows.map((row) => ({
id: row.id,
userId: row.user_id,
name: row.name,
timestamp: parseClickhouseUTCDateTimeFormat(row.timestamp),
}));
};
export const deleteTraces = async (projectId: string, traceIds: string[]) => {
const query = `
DELETE FROM traces
WHERE project_id = {projectId: String}
AND id IN ({traceIds: Array(String)});
`;
await commandClickhouse({
query: query,
params: {
projectId,
traceIds,
},
});
};
@@ -0,0 +1,128 @@
import { ObservationLevel, Trace } from "@prisma/client";
import { parseClickhouseUTCDateTimeFormat } from "./clickhouse";
import { TraceRecordReadType } from "./definitions";
import { TracesTableReturnType } from "./traces";
import Decimal from "decimal.js";
import { ScoreAggregate } from "../../features/scores";
import { convertDateToClickhouseDateTime } from "../clickhouse/client";
export const convertTraceDomainToClickhouse = (
trace: Trace,
): TraceRecordReadType => {
return {
id: trace.id,
timestamp: convertDateToClickhouseDateTime(trace.timestamp),
name: trace.name,
user_id: trace.userId,
metadata: trace.metadata as Record<string, string>,
release: trace.release,
version: trace.version,
project_id: trace.projectId,
public: trace.public,
bookmarked: trace.bookmarked,
tags: trace.tags,
input: trace.input as string,
output: trace.output as string,
session_id: trace.sessionId,
created_at: convertDateToClickhouseDateTime(trace.createdAt),
updated_at: convertDateToClickhouseDateTime(trace.updatedAt),
event_ts: convertDateToClickhouseDateTime(new Date()),
is_deleted: 0,
};
};
export const convertClickhouseToDomain = (
record: TraceRecordReadType,
): Trace => {
return {
id: record.id,
projectId: record.project_id,
name: record.name ?? null,
timestamp: parseClickhouseUTCDateTimeFormat(record.timestamp),
tags: record.tags,
bookmarked: record.bookmarked,
release: record.release ?? null,
version: record.version ?? null,
userId: record.user_id ?? null,
sessionId: record.session_id ?? null,
public: record.public,
input: record.input ?? null,
output: record.output ?? null,
metadata: record.metadata,
createdAt: parseClickhouseUTCDateTimeFormat(record.created_at),
updatedAt: parseClickhouseUTCDateTimeFormat(record.updated_at),
externalId: null,
};
};
export type TracesAllReturnType = {
id: string;
timestamp: Date;
name: string | null;
projectId: string;
userId: string | null;
release: string | null;
version: string | null;
public: boolean;
bookmarked: boolean;
sessionId: string | null;
tags: string[];
};
export const convertToReturnType = (
row: TracesTableReturnType,
): TracesAllReturnType => {
return {
id: row.id,
name: row.name ?? null,
timestamp: parseClickhouseUTCDateTimeFormat(row.timestamp),
tags: row.tags,
bookmarked: row.bookmarked,
release: row.release ?? null,
version: row.version ?? null,
projectId: row.project_id,
userId: row.user_id ?? null,
sessionId: row.session_id ?? null,
public: row.public,
};
};
export type TracesMetricsReturnType = {
id: string;
promptTokens: bigint;
completionTokens: bigint;
totalTokens: bigint;
latency: number | null;
level: ObservationLevel;
observationCount: bigint;
calculatedTotalCost: Decimal | null;
calculatedInputCost: Decimal | null;
calculatedOutputCost: Decimal | null;
scores: ScoreAggregate;
};
export const convertMetricsReturnType = (
row: TracesTableReturnType & { scores: ScoreAggregate },
): TracesMetricsReturnType => {
return {
id: row.id,
promptTokens: BigInt(row.usage_details?.input ?? 0),
completionTokens: BigInt(row.usage_details?.output ?? 0),
totalTokens: BigInt(row.usage_details?.total ?? 0),
latency: row.latency_milliseconds
? Number(row.latency_milliseconds) / 1000
: null,
level: row.level,
observationCount: BigInt(row.observation_count ?? 0),
calculatedTotalCost: row.cost_details?.total
? new Decimal(row.cost_details.total)
: null,
calculatedInputCost: row.cost_details?.input
? new Decimal(row.cost_details.input)
: null,
calculatedOutputCost: row.cost_details?.output
? new Decimal(row.cost_details.output)
: null,
scores: row.scores,
};
};
@@ -0,0 +1,3 @@
export type TableCount = {
count: number;
};
@@ -30,7 +30,7 @@ export class PromptService {
);
if (cachedPrompt) {
this.logInfo("Returning cached prompt for params", params);
this.logDebug("Returning cached prompt for params", params);
return cachedPrompt;
}
@@ -223,6 +223,10 @@ export class PromptService {
logger.info(`[PromptService] ${message}`, ...args);
}
private logDebug(message: string, ...args: any[]) {
logger.debug(`[PromptService] ${message}`, ...args);
}
private incrementMetric(name: Metrics, value: number = 1) {
try {
this.metricIncrementer?.(name, value);
@@ -1,6 +1,7 @@
import type { Readable } from "stream";
import {
GetObjectCommand,
ListObjectsV2Command,
PutObjectCommand,
S3Client,
} from "@aws-sdk/client-s3";
@@ -103,6 +104,23 @@ export class S3StorageService {
}
}
public async listFiles(prefix: string): Promise<string[]> {
const listCommand = new ListObjectsV2Command({
Bucket: this.bucketName,
Prefix: prefix,
});
try {
const response = await this.client.send(listCommand);
return (
response.Contents?.flatMap((file) => (file.Key ? [file.Key] : [])) ?? []
);
} catch (err) {
logger.error(`Failed to list files from S3 ${prefix}`, err);
throw Error("Failed to list files from S3");
}
}
private async getSignedUrl(
fileName: string,
ttlSeconds: number,
@@ -1,2 +1,6 @@
export * from "./sessionsView";
export * from "./types";
export * from "./mapObservationsTable";
export * from "./mapTracesTable";
export * from "./mapDashboards";
export * from "./mapScoresTable";
@@ -0,0 +1,76 @@
import { UiColumnMapping } from "./types";
export const dashboardColumnDefinitions: UiColumnMapping[] = [
{
uiTableName: "Trace Name",
uiTableId: "traceName",
clickhouseTableName: "traces",
clickhouseSelect: 't."name"',
},
{
uiTableName: "Tags",
uiTableId: "traceTags",
clickhouseTableName: "traces",
clickhouseSelect: 't."tags"',
},
{
uiTableName: "Timestamp",
uiTableId: "timestamp",
clickhouseTableName: "traces",
clickhouseSelect: 't."timestamp"',
},
{
clickhouseTableName: "scores",
clickhouseSelect: "timestamp",
uiTableId: "scoreTimestamp",
uiTableName: "Score Timestamp",
},
{
clickhouseTableName: "scores",
clickhouseSelect: "data_type",
uiTableId: "scoreDataType",
uiTableName: "Scores Data Type",
},
{
clickhouseTableName: "scores",
clickhouseSelect: "value",
uiTableId: "value",
uiTableName: "value",
},
{
clickhouseTableName: "observations",
clickhouseSelect: "o.start_time",
uiTableId: "startTime",
uiTableName: "Start Time",
},
{
clickhouseTableName: "observations",
clickhouseSelect: "o.end_time",
uiTableId: "endTime",
uiTableName: "End Time",
},
{
clickhouseTableName: "observations",
clickhouseSelect: "o.type",
uiTableId: "type",
uiTableName: "Type",
},
{
clickhouseTableName: "traces",
clickhouseSelect: "t.user_id",
uiTableId: "userId",
uiTableName: "User",
},
{
clickhouseTableName: "traces",
clickhouseSelect: "t.release",
uiTableId: "release",
uiTableName: "Release",
},
{
clickhouseTableName: "traces",
clickhouseSelect: "t.version",
uiTableId: "version",
uiTableName: "Version",
},
];
@@ -0,0 +1,184 @@
// This structure is maintained to relate the frontend table definitions with the clickhouse table definitions.
// The frontend only sends the column names to the backend. This needs to be changed in the future to send column IDs.
import { UiColumnMapping } from "./types";
export const observationsTableTraceUiColumnDefinitions: UiColumnMapping[] = [
{
uiTableName: "Trace Tags",
uiTableId: "traceTags",
clickhouseTableName: "traces",
clickhouseSelect: "t.tags",
},
{
uiTableName: "User ID",
uiTableId: "userId",
clickhouseTableName: "traces",
clickhouseSelect: 't."user_id"',
},
{
uiTableName: "Trace Name",
uiTableId: "traceName",
clickhouseTableName: "observations",
clickhouseSelect: 't."name"',
},
];
export const observationsTableUiColumnDefinitions: UiColumnMapping[] = [
...observationsTableTraceUiColumnDefinitions,
{
uiTableName: "ID",
uiTableId: "id",
clickhouseTableName: "observations",
clickhouseSelect: 'o."id"',
},
{
uiTableName: "Type",
uiTableId: "type",
clickhouseTableName: "observations",
clickhouseSelect: 'o."type"',
},
{
uiTableName: "Name",
uiTableId: "name",
clickhouseTableName: "observations",
clickhouseSelect: 'o."name"',
},
{
uiTableName: "Trace ID",
uiTableId: "traceId",
clickhouseTableName: "observations",
clickhouseSelect: 'o."trace_id"',
},
{
uiTableName: "Start Time",
uiTableId: "startTime",
clickhouseTableName: "observations",
clickhouseSelect: 'o."start_time"',
},
{
uiTableName: "End Time",
uiTableId: "endTime",
clickhouseTableName: "observations",
clickhouseSelect: 'o."end_time"',
},
{
uiTableName: "Time To First Token (s)",
uiTableId: "timeToFirstToken",
clickhouseTableName: "observations",
clickhouseSelect:
"if(isNull(completion_start_time), NULL, date_diff('seconds', start_time, completion_start_time))",
},
{
uiTableName: "Latency (s)",
uiTableId: "latency",
clickhouseTableName: "observations",
clickhouseSelect:
"if(isNull(end_time), NULL, date_diff('seconds', start_time, end_time))",
},
{
uiTableName: "Tokens per second",
uiTableId: "tokensPerSecond",
clickhouseTableName: "observations",
clickhouseSelect:
"usage_details['input'] / date_diff('seconds', start_time, end_time)",
},
{
uiTableName: "Input Cost ($)",
uiTableId: "inputCost",
clickhouseTableName: "observations",
clickhouseSelect:
"if(mapExists((k, v) -> (k = 'input'), cost_details), cost_details['input'], NULL)",
},
{
uiTableName: "Output Cost ($)",
uiTableId: "outputCost",
clickhouseTableName: "observations",
clickhouseSelect:
"if(mapExists((k, v) -> (k = 'output'), cost_details), cost_details['output'], NULL)",
},
{
uiTableName: "Total Cost ($)",
uiTableId: "totalCost",
clickhouseTableName: "observations",
clickhouseSelect:
"if(mapExists((k, v) -> (k = 'total'), cost_details), cost_details['total'], NULL)",
},
{
uiTableName: "Level",
uiTableId: "level",
clickhouseTableName: "observations",
clickhouseSelect: 'o."level"',
},
{
uiTableName: "Status Message",
uiTableId: "statusMessage",
clickhouseTableName: "observations",
clickhouseSelect: 'o."status_message"',
},
{
uiTableName: "Model",
uiTableId: "model",
clickhouseTableName: "observations",
clickhouseSelect: 'o."provided_model_name"',
},
{
uiTableName: "Input Tokens",
uiTableId: "inputTokens",
clickhouseTableName: "observations",
clickhouseSelect:
"if(mapExists((k, v) -> (k = 'input'), usage_details), usage_details['input'], NULL)",
},
{
uiTableName: "Output Tokens",
uiTableId: "outputTokens",
clickhouseTableName: "observations",
clickhouseSelect:
"if(mapExists((k, v) -> (k = 'output'), usage_details), usage_details['output'], NULL)",
},
{
uiTableName: "Total Tokens",
uiTableId: "totalTokens",
clickhouseTableName: "observations",
clickhouseSelect:
"if(mapExists((k, v) -> (k = 'total'), usage_details), usage_details['total'], NULL)",
},
{
uiTableName: "Usage",
uiTableId: "usage",
clickhouseTableName: "observations",
clickhouseSelect:
"if(mapExists((k, v) -> (k = 'total'), usage_details), usage_details['total'], NULL)",
},
{
uiTableName: "Metadata",
uiTableId: "metadata",
clickhouseTableName: "observations",
clickhouseSelect: 'o."metadata"',
},
{
uiTableName: "Scores",
uiTableId: "scores",
clickhouseTableName: "observations",
clickhouseSelect: "s_avg.scores_avg",
},
{
uiTableName: "Version",
uiTableId: "version",
clickhouseTableName: "observations",
clickhouseSelect: 'o."version"',
},
{
uiTableName: "Prompt Name",
uiTableId: "promptName",
clickhouseTableName: "observations",
clickhouseSelect: "o.prompt_name",
},
{
uiTableName: "Prompt Version",
uiTableId: "promptVersion",
clickhouseTableName: "observations",
clickhouseSelect: "o.prompt_version",
},
];

Some files were not shown because too many files have changed in this diff Show More