fix(docs): Add Example for Evaluator Extension - #3

Merged
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example
Feb 3, 2026
Merged

fix(docs): Add Example for Evaluator Extension#3
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example

Conversation

@namrataghadi-galileo

Copy link
Copy Markdown
Contributor

Deep Eval Evaluator example to show how Evaluator can be extended to create custom evaluators.

abhinav-galileoand others added 6 commits January 28, 2026 22:04
- Rename `plugins/` package to `evaluators/` with updated class names:
- RegexPlugin → RegexEvaluator
- ListPlugin → ListEvaluator
- JSONControlEvaluatorPlugin → JSONEvaluator
- SQLControlEvaluatorPlugin → SQLEvaluator
- Luna2Plugin → Luna2Evaluator
- Rename config classes to follow XEvaluatorConfig pattern:
- RegexConfig → RegexEvaluatorConfig
- ListConfig → ListEvaluatorConfig
- JSONControlEvaluatorPluginConfig → JSONEvaluatorConfig
- SQLControlEvaluatorPluginConfig → SQLEvaluatorConfig
- Luna2Config → Luna2EvaluatorConfig
- Update models package:
- Rename plugin.py to evaluator.py
- PluginMetadata → EvaluatorMetadata
- PluginEvaluator → Evaluator
- register_plugin → register_evaluator
- EvaluatorConfig.plugin field → EvaluatorConfig.name
- Update engine package:
- discover_plugins → discover_evaluators
- list_plugins → list_evaluators
- Entry point: agent_control.plugins → agent_control.evaluators
- get_evaluator → get_evaluator_instance (to avoid name collision)
- Update server:
- /api/v1/plugins endpoint → /api/v1/evaluators
- Add migration to rename plugin column to evaluator in evaluator_configs table
- Remove all backwards compatibility aliases
- Update all documentation, examples, and tests
@codecov

codecovBot commented Jan 31, 2026

Copy link
Copy Markdown

The author of this PR, namrataghadi-galileo, is not an activated member of this organization on Codecov.
Please activate this user on Codecov to display this PR comment.
Coverage data is still being uploaded to Codecov.io for purposes of overall coverage calculations.
Please don't hesitate to email us at support@codecov.io with any questions.

Comment threadexamples/deepeval/evaluator.py Outdated
Comment threadexamples/deepeval/pyproject.toml Outdated
@namrataghadi-galileo
namrataghadi-galileo merged commit c2a70b3 into mainFeb 3, 2026
5 checks passed
@namrataghadi-galileo
namrataghadi-galileo deleted the feature/add-deepeval-evaluator-example branch February 3, 2026 22:19
galileo-automation pushed a commit that referenced this pull request Mar 4, 2026
## 1.0.0 (2026-03-04)
### ⚠ BREAKING CHANGES
* **server:** Feature/56688 fix image bug (#48)
* **sdk:** a bug in docker file (#46)
* **server:** Feature/56688 fix docker and create bash (#45)
* **evaluators:** Evaluator reorganization with new package structure
Package Structure:
- agent-control-evaluators (v3.0.0): core + regex, list, json, sql
- agent-control-evaluator-galileo (v3.0.0): Luna2 evaluator
Key Changes:
- Entry points for evaluator discovery (agent_control.evaluators)
- Dot notation for external evaluators (galileo.luna2 not galileo/luna2)
- Dynamic __version__ via importlib.metadata
- Server uses evaluators as runtime dep (no longer vendored)
- Release workflow publishes both packages to PyPI
Bug Fixes:
- JSON evaluator: field_constraints/field_patterns in extra-fields allow-list
- SQL evaluator: LIMIT/OFFSET bypass fix
Migration:
- Import: agent_control_evaluator_galileo.luna2 (not agent_control_evaluators.galileo_luna2)
- DB: UPDATE controls SET evaluator.name replace('/', '.')
* **server:** add time-series stats and split API endpoints (#6)
* **evaluators:** rename plugin to evaluator throughout (#81)
* **models:** simplify step model and schema (#70)
### Features
* Add plugin auto-discovery via Python entry points ([#49](#49)) ([1521182](1521182))
* **docs:** add GitHub badges and CI coverage reporting ([#90](#90)) ([be1fa14](be1fa14))
* **evaluators:** add required_column_values for multi-tenant SQL validation ([#30](#30)) ([532386c](532386c))
* **sdk-ts:** automate semantic-release for npm publishing ([#52](#52)) ([2b43958](2b43958))
* **sdk:** Add PyPI packaging with semantic release ([#52](#52)) ([7c24f7f](7c24f7f))
* **sdk:** Auto-populate init() steps from [@control](https://github.com/control)() decorators ([#23](#23)) ([dc0f2a4](dc0f2a4))
* **sdk:** export ControlScope, ControlMatch, and EvaluatorResult models ([#18](#18)) ([0d49cad](0d49cad))
* **sdk:** Get Agent Controls from SDK Init ([#15](#15)) ([a485f93](a485f93))
* **sdk:** Refresh controls in a background loop ([#43](#43)) ([03f826d](03f826d))
* **sdk:** ship TypeScript SDK with deterministic method naming ([#32](#32)) ([a76e9b0](a76e9b0))
* **server:** add evaluator config store ([#78](#78)) ([cc14aa6](cc14aa6))
* **server:** add initAgent conflict_mode overwrite mode with SDK defaults ([#40](#40)) ([f3ed2b8](f3ed2b8))
* **server:** Add observability system for control execution tracking ([#44](#44)) ([fd0bddc](fd0bddc))
* **server:** add prometheus metrics for endpoints ([#68](#68)) ([775612c](775612c))
* **server:** add time-series stats and split API endpoints ([#6](#6)) ([a0fa597](a0fa597))
* **server:** hard-cut migrate to remove agent UUID ([#44](#44)) ([ee322c9](ee322c9))
* **server:** Optional Policy and many to many relationships ([#41](#41)) ([1a62746](1a62746))
* **ui:** add sql, luna2, json control forms and restructure the code ([#54](#54)) ([c4c1d4a](c4c1d4a))
* **ui:** allow to delete control ([#39](#39)) ([7dc4ca3](7dc4ca3))
* **ui:** Control Store Flow Updated ([#4](#4)) ([dda9f70](dda9f70))
* **ui:** stats dashboard ([#80](#80)) ([4cbb7fe](4cbb7fe))
* **ui:** Steps dropdown rendered based on api return values ([#36](#36)) ([a2aca43](a2aca43))
* **ui:** tests added and some minor ui changes, added error boundaries ([#61](#61)) ([009852b](009852b))
* **ui:** update agent control icon and favicon ([#42](#42)) ([19af8fa](19af8fa))
### Bug Fixes
* **ci:** Add ui scope to PR title validation ([#59](#59)) ([e0fdb52](e0fdb52))
* **ci:** correct galileo contrib path in release build script ([#51](#51)) ([2de6013](2de6013))
* **ci:** Enable pr title on prs ([#56](#56)) ([3d8b5fe](3d8b5fe))
* **ci:** Fix release ([#11](#11)) ([9dd3dd7](9dd3dd7))
* **ci:** Use galileo-automation bot for releases ([#57](#57)) ([bc8eea0](bc8eea0))
* **docs:** Add Example for Evaluator Extension ([#3](#3)) ([c2a70b3](c2a70b3))
* **docs:** add setup script ([#49](#49)) ([7a212c3](7a212c3))
* **docs:** Clean up Protect ([#76](#76)) ([99c16fd](99c16fd))
* **docs:** Fix Examples for LangGraph ([#64](#64)) ([23b30ae](23b30ae))
* **docs:** Improve documentation for open source release ([#47](#47)) ([9018fb3](9018fb3))
* **docs:** Remove old/unused examples ([#66](#66)) ([f417781](f417781))
* **docs:** Update Contributing Guide ([#8](#8)) ([10b34c8](10b34c8))
* **docs:** Update readme ([#37](#37)) ([7531d83](7531d83))
* **docs:** Update README ([#2](#2)) ([379bb15](379bb15))
* **examples:** Control sets cleanup with signed ([#65](#65)) ([af7b5fb](af7b5fb))
* **examples:** Update crew ai example to use evaluator ([#93](#93)) ([1c65084](1c65084))
* **infra:** Add plugins directory to Dockerfile ([#58](#58)) ([171d459](171d459))
* **infra:** install engine/evaluators in server image ([#14](#14)) ([d5ae157](d5ae157))
* **models:** use StrEnum for error enums ([#12](#12)) ([3f41c9f](3f41c9f))
* **sdk-ts:** add conventional commits preset dependency ([#55](#55)) ([540fe9d](540fe9d))
* **sdk-ts:** export npm token for semantic-release npm auth ([#54](#54)) ([1b6b993](1b6b993))
* **sdk:** 54253 add steer action and example ([#38](#38)) ([bf2380a](bf2380a))
* **sdk:** a bug in docker file ([#46](#46)) ([12d1794](12d1794))
* **sdk:** Add step_name as parameter to control ([#25](#25)) ([19ade9d](19ade9d))
* **sdk:** emit observability events for SDK-evaluated controls and fix non_matches propagation ([#24](#24)) ([6a9da69](6a9da69))
* **sdk:** enforce UUID agent IDs ([#9](#9)) ([5ccdbd0](5ccdbd0))
* **sdk:** Fix logging ([#77](#77)) ([b1f078c](b1f078c))
* **sdk:** plugin to evaluator.. agent_protect to agent_control ([#88](#88)) ([fc9b088](fc9b088))
* **server:** enforce public-safe API error responses ([#20](#20)) ([e50d817](e50d817))
* **server:** Feature/56688 fix docker and create bash ([#45](#45)) ([7277e27](7277e27))
* **server:** Feature/56688 fix image bug ([#48](#48)) ([71e6b44](71e6b44))
* **server:** fix alembic migrations ([#47](#47)) ([c19c17c](c19c17c))
* **server:** reject initAgent UUID/name mismatch ([#13](#13)) ([19d61ff](19d61ff))
* tighten evaluation error handling and preserve control data ([52a1ef8](52a1ef8))
* **ui:** Fix UI and clients for simplified step schema ([#75](#75)) ([be2aaf0](be2aaf0))
* **ui:** json validation ([#10](#10)) ([a0cd5af](a0cd5af))
* **ui:** selector subpaths issue ([#34](#34)) ([79cb776](79cb776))
* **ui:** UI feedback fixes ([#27](#27)) ([6004761](6004761))
### Code Refactoring
* **evaluators:** rename plugin to evaluator throughout ([#81](#81)) ([0134682](0134682))
* **evaluators:** split into builtin + extra packages for PyPI ([#5](#5)) ([0e0a78a](0e0a78a))
* **models:** simplify step model and schema ([#70](#70)) ([4c1d637](4c1d637))
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Mar 21, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Jun 12, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@namrataghadi-galileo@lan17@nachiket-galileo@abhinav-galileo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(docs): Add Example for Evaluator Extension - #3

Merged
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example
Feb 3, 2026
Merged

fix(docs): Add Example for Evaluator Extension#3
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example

Conversation

@namrataghadi-galileo

Copy link
Copy Markdown
Contributor

Deep Eval Evaluator example to show how Evaluator can be extended to create custom evaluators.

abhinav-galileoand others added 6 commits January 28, 2026 22:04
- Rename `plugins/` package to `evaluators/` with updated class names:
- RegexPlugin → RegexEvaluator
- ListPlugin → ListEvaluator
- JSONControlEvaluatorPlugin → JSONEvaluator
- SQLControlEvaluatorPlugin → SQLEvaluator
- Luna2Plugin → Luna2Evaluator
- Rename config classes to follow XEvaluatorConfig pattern:
- RegexConfig → RegexEvaluatorConfig
- ListConfig → ListEvaluatorConfig
- JSONControlEvaluatorPluginConfig → JSONEvaluatorConfig
- SQLControlEvaluatorPluginConfig → SQLEvaluatorConfig
- Luna2Config → Luna2EvaluatorConfig
- Update models package:
- Rename plugin.py to evaluator.py
- PluginMetadata → EvaluatorMetadata
- PluginEvaluator → Evaluator
- register_plugin → register_evaluator
- EvaluatorConfig.plugin field → EvaluatorConfig.name
- Update engine package:
- discover_plugins → discover_evaluators
- list_plugins → list_evaluators
- Entry point: agent_control.plugins → agent_control.evaluators
- get_evaluator → get_evaluator_instance (to avoid name collision)
- Update server:
- /api/v1/plugins endpoint → /api/v1/evaluators
- Add migration to rename plugin column to evaluator in evaluator_configs table
- Remove all backwards compatibility aliases
- Update all documentation, examples, and tests
@codecov

codecovBot commented Jan 31, 2026

Copy link
Copy Markdown

The author of this PR, namrataghadi-galileo, is not an activated member of this organization on Codecov.
Please activate this user on Codecov to display this PR comment.
Coverage data is still being uploaded to Codecov.io for purposes of overall coverage calculations.
Please don't hesitate to email us at support@codecov.io with any questions.

Comment threadexamples/deepeval/evaluator.py Outdated
Comment threadexamples/deepeval/pyproject.toml Outdated
@namrataghadi-galileo
namrataghadi-galileo merged commit c2a70b3 into mainFeb 3, 2026
5 checks passed
@namrataghadi-galileo
namrataghadi-galileo deleted the feature/add-deepeval-evaluator-example branch February 3, 2026 22:19
galileo-automation pushed a commit that referenced this pull request Mar 4, 2026
## 1.0.0 (2026-03-04)
### ⚠ BREAKING CHANGES
* **server:** Feature/56688 fix image bug (#48)
* **sdk:** a bug in docker file (#46)
* **server:** Feature/56688 fix docker and create bash (#45)
* **evaluators:** Evaluator reorganization with new package structure
Package Structure:
- agent-control-evaluators (v3.0.0): core + regex, list, json, sql
- agent-control-evaluator-galileo (v3.0.0): Luna2 evaluator
Key Changes:
- Entry points for evaluator discovery (agent_control.evaluators)
- Dot notation for external evaluators (galileo.luna2 not galileo/luna2)
- Dynamic __version__ via importlib.metadata
- Server uses evaluators as runtime dep (no longer vendored)
- Release workflow publishes both packages to PyPI
Bug Fixes:
- JSON evaluator: field_constraints/field_patterns in extra-fields allow-list
- SQL evaluator: LIMIT/OFFSET bypass fix
Migration:
- Import: agent_control_evaluator_galileo.luna2 (not agent_control_evaluators.galileo_luna2)
- DB: UPDATE controls SET evaluator.name replace('/', '.')
* **server:** add time-series stats and split API endpoints (#6)
* **evaluators:** rename plugin to evaluator throughout (#81)
* **models:** simplify step model and schema (#70)
### Features
* Add plugin auto-discovery via Python entry points ([#49](#49)) ([1521182](1521182))
* **docs:** add GitHub badges and CI coverage reporting ([#90](#90)) ([be1fa14](be1fa14))
* **evaluators:** add required_column_values for multi-tenant SQL validation ([#30](#30)) ([532386c](532386c))
* **sdk-ts:** automate semantic-release for npm publishing ([#52](#52)) ([2b43958](2b43958))
* **sdk:** Add PyPI packaging with semantic release ([#52](#52)) ([7c24f7f](7c24f7f))
* **sdk:** Auto-populate init() steps from [@control](https://github.com/control)() decorators ([#23](#23)) ([dc0f2a4](dc0f2a4))
* **sdk:** export ControlScope, ControlMatch, and EvaluatorResult models ([#18](#18)) ([0d49cad](0d49cad))
* **sdk:** Get Agent Controls from SDK Init ([#15](#15)) ([a485f93](a485f93))
* **sdk:** Refresh controls in a background loop ([#43](#43)) ([03f826d](03f826d))
* **sdk:** ship TypeScript SDK with deterministic method naming ([#32](#32)) ([a76e9b0](a76e9b0))
* **server:** add evaluator config store ([#78](#78)) ([cc14aa6](cc14aa6))
* **server:** add initAgent conflict_mode overwrite mode with SDK defaults ([#40](#40)) ([f3ed2b8](f3ed2b8))
* **server:** Add observability system for control execution tracking ([#44](#44)) ([fd0bddc](fd0bddc))
* **server:** add prometheus metrics for endpoints ([#68](#68)) ([775612c](775612c))
* **server:** add time-series stats and split API endpoints ([#6](#6)) ([a0fa597](a0fa597))
* **server:** hard-cut migrate to remove agent UUID ([#44](#44)) ([ee322c9](ee322c9))
* **server:** Optional Policy and many to many relationships ([#41](#41)) ([1a62746](1a62746))
* **ui:** add sql, luna2, json control forms and restructure the code ([#54](#54)) ([c4c1d4a](c4c1d4a))
* **ui:** allow to delete control ([#39](#39)) ([7dc4ca3](7dc4ca3))
* **ui:** Control Store Flow Updated ([#4](#4)) ([dda9f70](dda9f70))
* **ui:** stats dashboard ([#80](#80)) ([4cbb7fe](4cbb7fe))
* **ui:** Steps dropdown rendered based on api return values ([#36](#36)) ([a2aca43](a2aca43))
* **ui:** tests added and some minor ui changes, added error boundaries ([#61](#61)) ([009852b](009852b))
* **ui:** update agent control icon and favicon ([#42](#42)) ([19af8fa](19af8fa))
### Bug Fixes
* **ci:** Add ui scope to PR title validation ([#59](#59)) ([e0fdb52](e0fdb52))
* **ci:** correct galileo contrib path in release build script ([#51](#51)) ([2de6013](2de6013))
* **ci:** Enable pr title on prs ([#56](#56)) ([3d8b5fe](3d8b5fe))
* **ci:** Fix release ([#11](#11)) ([9dd3dd7](9dd3dd7))
* **ci:** Use galileo-automation bot for releases ([#57](#57)) ([bc8eea0](bc8eea0))
* **docs:** Add Example for Evaluator Extension ([#3](#3)) ([c2a70b3](c2a70b3))
* **docs:** add setup script ([#49](#49)) ([7a212c3](7a212c3))
* **docs:** Clean up Protect ([#76](#76)) ([99c16fd](99c16fd))
* **docs:** Fix Examples for LangGraph ([#64](#64)) ([23b30ae](23b30ae))
* **docs:** Improve documentation for open source release ([#47](#47)) ([9018fb3](9018fb3))
* **docs:** Remove old/unused examples ([#66](#66)) ([f417781](f417781))
* **docs:** Update Contributing Guide ([#8](#8)) ([10b34c8](10b34c8))
* **docs:** Update readme ([#37](#37)) ([7531d83](7531d83))
* **docs:** Update README ([#2](#2)) ([379bb15](379bb15))
* **examples:** Control sets cleanup with signed ([#65](#65)) ([af7b5fb](af7b5fb))
* **examples:** Update crew ai example to use evaluator ([#93](#93)) ([1c65084](1c65084))
* **infra:** Add plugins directory to Dockerfile ([#58](#58)) ([171d459](171d459))
* **infra:** install engine/evaluators in server image ([#14](#14)) ([d5ae157](d5ae157))
* **models:** use StrEnum for error enums ([#12](#12)) ([3f41c9f](3f41c9f))
* **sdk-ts:** add conventional commits preset dependency ([#55](#55)) ([540fe9d](540fe9d))
* **sdk-ts:** export npm token for semantic-release npm auth ([#54](#54)) ([1b6b993](1b6b993))
* **sdk:** 54253 add steer action and example ([#38](#38)) ([bf2380a](bf2380a))
* **sdk:** a bug in docker file ([#46](#46)) ([12d1794](12d1794))
* **sdk:** Add step_name as parameter to control ([#25](#25)) ([19ade9d](19ade9d))
* **sdk:** emit observability events for SDK-evaluated controls and fix non_matches propagation ([#24](#24)) ([6a9da69](6a9da69))
* **sdk:** enforce UUID agent IDs ([#9](#9)) ([5ccdbd0](5ccdbd0))
* **sdk:** Fix logging ([#77](#77)) ([b1f078c](b1f078c))
* **sdk:** plugin to evaluator.. agent_protect to agent_control ([#88](#88)) ([fc9b088](fc9b088))
* **server:** enforce public-safe API error responses ([#20](#20)) ([e50d817](e50d817))
* **server:** Feature/56688 fix docker and create bash ([#45](#45)) ([7277e27](7277e27))
* **server:** Feature/56688 fix image bug ([#48](#48)) ([71e6b44](71e6b44))
* **server:** fix alembic migrations ([#47](#47)) ([c19c17c](c19c17c))
* **server:** reject initAgent UUID/name mismatch ([#13](#13)) ([19d61ff](19d61ff))
* tighten evaluation error handling and preserve control data ([52a1ef8](52a1ef8))
* **ui:** Fix UI and clients for simplified step schema ([#75](#75)) ([be2aaf0](be2aaf0))
* **ui:** json validation ([#10](#10)) ([a0cd5af](a0cd5af))
* **ui:** selector subpaths issue ([#34](#34)) ([79cb776](79cb776))
* **ui:** UI feedback fixes ([#27](#27)) ([6004761](6004761))
### Code Refactoring
* **evaluators:** rename plugin to evaluator throughout ([#81](#81)) ([0134682](0134682))
* **evaluators:** split into builtin + extra packages for PyPI ([#5](#5)) ([0e0a78a](0e0a78a))
* **models:** simplify step model and schema ([#70](#70)) ([4c1d637](4c1d637))
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Mar 21, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Jun 12, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@namrataghadi-galileo@lan17@nachiket-galileo@abhinav-galileo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(docs): Add Example for Evaluator Extension - #3

Merged
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example
Feb 3, 2026
Merged

fix(docs): Add Example for Evaluator Extension#3
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example

Conversation

@namrataghadi-galileo

Copy link
Copy Markdown
Contributor

Deep Eval Evaluator example to show how Evaluator can be extended to create custom evaluators.

abhinav-galileoand others added 6 commits January 28, 2026 22:04
- Rename `plugins/` package to `evaluators/` with updated class names:
- RegexPlugin → RegexEvaluator
- ListPlugin → ListEvaluator
- JSONControlEvaluatorPlugin → JSONEvaluator
- SQLControlEvaluatorPlugin → SQLEvaluator
- Luna2Plugin → Luna2Evaluator
- Rename config classes to follow XEvaluatorConfig pattern:
- RegexConfig → RegexEvaluatorConfig
- ListConfig → ListEvaluatorConfig
- JSONControlEvaluatorPluginConfig → JSONEvaluatorConfig
- SQLControlEvaluatorPluginConfig → SQLEvaluatorConfig
- Luna2Config → Luna2EvaluatorConfig
- Update models package:
- Rename plugin.py to evaluator.py
- PluginMetadata → EvaluatorMetadata
- PluginEvaluator → Evaluator
- register_plugin → register_evaluator
- EvaluatorConfig.plugin field → EvaluatorConfig.name
- Update engine package:
- discover_plugins → discover_evaluators
- list_plugins → list_evaluators
- Entry point: agent_control.plugins → agent_control.evaluators
- get_evaluator → get_evaluator_instance (to avoid name collision)
- Update server:
- /api/v1/plugins endpoint → /api/v1/evaluators
- Add migration to rename plugin column to evaluator in evaluator_configs table
- Remove all backwards compatibility aliases
- Update all documentation, examples, and tests
@codecov

codecovBot commented Jan 31, 2026

Copy link
Copy Markdown

The author of this PR, namrataghadi-galileo, is not an activated member of this organization on Codecov.
Please activate this user on Codecov to display this PR comment.
Coverage data is still being uploaded to Codecov.io for purposes of overall coverage calculations.
Please don't hesitate to email us at support@codecov.io with any questions.

Comment threadexamples/deepeval/evaluator.py Outdated
Comment threadexamples/deepeval/pyproject.toml Outdated
@namrataghadi-galileo
namrataghadi-galileo merged commit c2a70b3 into mainFeb 3, 2026
5 checks passed
@namrataghadi-galileo
namrataghadi-galileo deleted the feature/add-deepeval-evaluator-example branch February 3, 2026 22:19
galileo-automation pushed a commit that referenced this pull request Mar 4, 2026
## 1.0.0 (2026-03-04)
### ⚠ BREAKING CHANGES
* **server:** Feature/56688 fix image bug (#48)
* **sdk:** a bug in docker file (#46)
* **server:** Feature/56688 fix docker and create bash (#45)
* **evaluators:** Evaluator reorganization with new package structure
Package Structure:
- agent-control-evaluators (v3.0.0): core + regex, list, json, sql
- agent-control-evaluator-galileo (v3.0.0): Luna2 evaluator
Key Changes:
- Entry points for evaluator discovery (agent_control.evaluators)
- Dot notation for external evaluators (galileo.luna2 not galileo/luna2)
- Dynamic __version__ via importlib.metadata
- Server uses evaluators as runtime dep (no longer vendored)
- Release workflow publishes both packages to PyPI
Bug Fixes:
- JSON evaluator: field_constraints/field_patterns in extra-fields allow-list
- SQL evaluator: LIMIT/OFFSET bypass fix
Migration:
- Import: agent_control_evaluator_galileo.luna2 (not agent_control_evaluators.galileo_luna2)
- DB: UPDATE controls SET evaluator.name replace('/', '.')
* **server:** add time-series stats and split API endpoints (#6)
* **evaluators:** rename plugin to evaluator throughout (#81)
* **models:** simplify step model and schema (#70)
### Features
* Add plugin auto-discovery via Python entry points ([#49](#49)) ([1521182](1521182))
* **docs:** add GitHub badges and CI coverage reporting ([#90](#90)) ([be1fa14](be1fa14))
* **evaluators:** add required_column_values for multi-tenant SQL validation ([#30](#30)) ([532386c](532386c))
* **sdk-ts:** automate semantic-release for npm publishing ([#52](#52)) ([2b43958](2b43958))
* **sdk:** Add PyPI packaging with semantic release ([#52](#52)) ([7c24f7f](7c24f7f))
* **sdk:** Auto-populate init() steps from [@control](https://github.com/control)() decorators ([#23](#23)) ([dc0f2a4](dc0f2a4))
* **sdk:** export ControlScope, ControlMatch, and EvaluatorResult models ([#18](#18)) ([0d49cad](0d49cad))
* **sdk:** Get Agent Controls from SDK Init ([#15](#15)) ([a485f93](a485f93))
* **sdk:** Refresh controls in a background loop ([#43](#43)) ([03f826d](03f826d))
* **sdk:** ship TypeScript SDK with deterministic method naming ([#32](#32)) ([a76e9b0](a76e9b0))
* **server:** add evaluator config store ([#78](#78)) ([cc14aa6](cc14aa6))
* **server:** add initAgent conflict_mode overwrite mode with SDK defaults ([#40](#40)) ([f3ed2b8](f3ed2b8))
* **server:** Add observability system for control execution tracking ([#44](#44)) ([fd0bddc](fd0bddc))
* **server:** add prometheus metrics for endpoints ([#68](#68)) ([775612c](775612c))
* **server:** add time-series stats and split API endpoints ([#6](#6)) ([a0fa597](a0fa597))
* **server:** hard-cut migrate to remove agent UUID ([#44](#44)) ([ee322c9](ee322c9))
* **server:** Optional Policy and many to many relationships ([#41](#41)) ([1a62746](1a62746))
* **ui:** add sql, luna2, json control forms and restructure the code ([#54](#54)) ([c4c1d4a](c4c1d4a))
* **ui:** allow to delete control ([#39](#39)) ([7dc4ca3](7dc4ca3))
* **ui:** Control Store Flow Updated ([#4](#4)) ([dda9f70](dda9f70))
* **ui:** stats dashboard ([#80](#80)) ([4cbb7fe](4cbb7fe))
* **ui:** Steps dropdown rendered based on api return values ([#36](#36)) ([a2aca43](a2aca43))
* **ui:** tests added and some minor ui changes, added error boundaries ([#61](#61)) ([009852b](009852b))
* **ui:** update agent control icon and favicon ([#42](#42)) ([19af8fa](19af8fa))
### Bug Fixes
* **ci:** Add ui scope to PR title validation ([#59](#59)) ([e0fdb52](e0fdb52))
* **ci:** correct galileo contrib path in release build script ([#51](#51)) ([2de6013](2de6013))
* **ci:** Enable pr title on prs ([#56](#56)) ([3d8b5fe](3d8b5fe))
* **ci:** Fix release ([#11](#11)) ([9dd3dd7](9dd3dd7))
* **ci:** Use galileo-automation bot for releases ([#57](#57)) ([bc8eea0](bc8eea0))
* **docs:** Add Example for Evaluator Extension ([#3](#3)) ([c2a70b3](c2a70b3))
* **docs:** add setup script ([#49](#49)) ([7a212c3](7a212c3))
* **docs:** Clean up Protect ([#76](#76)) ([99c16fd](99c16fd))
* **docs:** Fix Examples for LangGraph ([#64](#64)) ([23b30ae](23b30ae))
* **docs:** Improve documentation for open source release ([#47](#47)) ([9018fb3](9018fb3))
* **docs:** Remove old/unused examples ([#66](#66)) ([f417781](f417781))
* **docs:** Update Contributing Guide ([#8](#8)) ([10b34c8](10b34c8))
* **docs:** Update readme ([#37](#37)) ([7531d83](7531d83))
* **docs:** Update README ([#2](#2)) ([379bb15](379bb15))
* **examples:** Control sets cleanup with signed ([#65](#65)) ([af7b5fb](af7b5fb))
* **examples:** Update crew ai example to use evaluator ([#93](#93)) ([1c65084](1c65084))
* **infra:** Add plugins directory to Dockerfile ([#58](#58)) ([171d459](171d459))
* **infra:** install engine/evaluators in server image ([#14](#14)) ([d5ae157](d5ae157))
* **models:** use StrEnum for error enums ([#12](#12)) ([3f41c9f](3f41c9f))
* **sdk-ts:** add conventional commits preset dependency ([#55](#55)) ([540fe9d](540fe9d))
* **sdk-ts:** export npm token for semantic-release npm auth ([#54](#54)) ([1b6b993](1b6b993))
* **sdk:** 54253 add steer action and example ([#38](#38)) ([bf2380a](bf2380a))
* **sdk:** a bug in docker file ([#46](#46)) ([12d1794](12d1794))
* **sdk:** Add step_name as parameter to control ([#25](#25)) ([19ade9d](19ade9d))
* **sdk:** emit observability events for SDK-evaluated controls and fix non_matches propagation ([#24](#24)) ([6a9da69](6a9da69))
* **sdk:** enforce UUID agent IDs ([#9](#9)) ([5ccdbd0](5ccdbd0))
* **sdk:** Fix logging ([#77](#77)) ([b1f078c](b1f078c))
* **sdk:** plugin to evaluator.. agent_protect to agent_control ([#88](#88)) ([fc9b088](fc9b088))
* **server:** enforce public-safe API error responses ([#20](#20)) ([e50d817](e50d817))
* **server:** Feature/56688 fix docker and create bash ([#45](#45)) ([7277e27](7277e27))
* **server:** Feature/56688 fix image bug ([#48](#48)) ([71e6b44](71e6b44))
* **server:** fix alembic migrations ([#47](#47)) ([c19c17c](c19c17c))
* **server:** reject initAgent UUID/name mismatch ([#13](#13)) ([19d61ff](19d61ff))
* tighten evaluation error handling and preserve control data ([52a1ef8](52a1ef8))
* **ui:** Fix UI and clients for simplified step schema ([#75](#75)) ([be2aaf0](be2aaf0))
* **ui:** json validation ([#10](#10)) ([a0cd5af](a0cd5af))
* **ui:** selector subpaths issue ([#34](#34)) ([79cb776](79cb776))
* **ui:** UI feedback fixes ([#27](#27)) ([6004761](6004761))
### Code Refactoring
* **evaluators:** rename plugin to evaluator throughout ([#81](#81)) ([0134682](0134682))
* **evaluators:** split into builtin + extra packages for PyPI ([#5](#5)) ([0e0a78a](0e0a78a))
* **models:** simplify step model and schema ([#70](#70)) ([4c1d637](4c1d637))
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Mar 21, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Jun 12, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@namrataghadi-galileo@lan17@nachiket-galileo@abhinav-galileo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(docs): Add Example for Evaluator Extension - #3

Merged
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example
Feb 3, 2026
Merged

fix(docs): Add Example for Evaluator Extension#3
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example

Conversation

@namrataghadi-galileo

Copy link
Copy Markdown
Contributor

Deep Eval Evaluator example to show how Evaluator can be extended to create custom evaluators.

abhinav-galileoand others added 6 commits January 28, 2026 22:04
- Rename `plugins/` package to `evaluators/` with updated class names:
- RegexPlugin → RegexEvaluator
- ListPlugin → ListEvaluator
- JSONControlEvaluatorPlugin → JSONEvaluator
- SQLControlEvaluatorPlugin → SQLEvaluator
- Luna2Plugin → Luna2Evaluator
- Rename config classes to follow XEvaluatorConfig pattern:
- RegexConfig → RegexEvaluatorConfig
- ListConfig → ListEvaluatorConfig
- JSONControlEvaluatorPluginConfig → JSONEvaluatorConfig
- SQLControlEvaluatorPluginConfig → SQLEvaluatorConfig
- Luna2Config → Luna2EvaluatorConfig
- Update models package:
- Rename plugin.py to evaluator.py
- PluginMetadata → EvaluatorMetadata
- PluginEvaluator → Evaluator
- register_plugin → register_evaluator
- EvaluatorConfig.plugin field → EvaluatorConfig.name
- Update engine package:
- discover_plugins → discover_evaluators
- list_plugins → list_evaluators
- Entry point: agent_control.plugins → agent_control.evaluators
- get_evaluator → get_evaluator_instance (to avoid name collision)
- Update server:
- /api/v1/plugins endpoint → /api/v1/evaluators
- Add migration to rename plugin column to evaluator in evaluator_configs table
- Remove all backwards compatibility aliases
- Update all documentation, examples, and tests
@codecov

codecovBot commented Jan 31, 2026

Copy link
Copy Markdown

The author of this PR, namrataghadi-galileo, is not an activated member of this organization on Codecov.
Please activate this user on Codecov to display this PR comment.
Coverage data is still being uploaded to Codecov.io for purposes of overall coverage calculations.
Please don't hesitate to email us at support@codecov.io with any questions.

Comment threadexamples/deepeval/evaluator.py Outdated
Comment threadexamples/deepeval/pyproject.toml Outdated
@namrataghadi-galileo
namrataghadi-galileo merged commit c2a70b3 into mainFeb 3, 2026
5 checks passed
@namrataghadi-galileo
namrataghadi-galileo deleted the feature/add-deepeval-evaluator-example branch February 3, 2026 22:19
galileo-automation pushed a commit that referenced this pull request Mar 4, 2026
## 1.0.0 (2026-03-04)
### ⚠ BREAKING CHANGES
* **server:** Feature/56688 fix image bug (#48)
* **sdk:** a bug in docker file (#46)
* **server:** Feature/56688 fix docker and create bash (#45)
* **evaluators:** Evaluator reorganization with new package structure
Package Structure:
- agent-control-evaluators (v3.0.0): core + regex, list, json, sql
- agent-control-evaluator-galileo (v3.0.0): Luna2 evaluator
Key Changes:
- Entry points for evaluator discovery (agent_control.evaluators)
- Dot notation for external evaluators (galileo.luna2 not galileo/luna2)
- Dynamic __version__ via importlib.metadata
- Server uses evaluators as runtime dep (no longer vendored)
- Release workflow publishes both packages to PyPI
Bug Fixes:
- JSON evaluator: field_constraints/field_patterns in extra-fields allow-list
- SQL evaluator: LIMIT/OFFSET bypass fix
Migration:
- Import: agent_control_evaluator_galileo.luna2 (not agent_control_evaluators.galileo_luna2)
- DB: UPDATE controls SET evaluator.name replace('/', '.')
* **server:** add time-series stats and split API endpoints (#6)
* **evaluators:** rename plugin to evaluator throughout (#81)
* **models:** simplify step model and schema (#70)
### Features
* Add plugin auto-discovery via Python entry points ([#49](#49)) ([1521182](1521182))
* **docs:** add GitHub badges and CI coverage reporting ([#90](#90)) ([be1fa14](be1fa14))
* **evaluators:** add required_column_values for multi-tenant SQL validation ([#30](#30)) ([532386c](532386c))
* **sdk-ts:** automate semantic-release for npm publishing ([#52](#52)) ([2b43958](2b43958))
* **sdk:** Add PyPI packaging with semantic release ([#52](#52)) ([7c24f7f](7c24f7f))
* **sdk:** Auto-populate init() steps from [@control](https://github.com/control)() decorators ([#23](#23)) ([dc0f2a4](dc0f2a4))
* **sdk:** export ControlScope, ControlMatch, and EvaluatorResult models ([#18](#18)) ([0d49cad](0d49cad))
* **sdk:** Get Agent Controls from SDK Init ([#15](#15)) ([a485f93](a485f93))
* **sdk:** Refresh controls in a background loop ([#43](#43)) ([03f826d](03f826d))
* **sdk:** ship TypeScript SDK with deterministic method naming ([#32](#32)) ([a76e9b0](a76e9b0))
* **server:** add evaluator config store ([#78](#78)) ([cc14aa6](cc14aa6))
* **server:** add initAgent conflict_mode overwrite mode with SDK defaults ([#40](#40)) ([f3ed2b8](f3ed2b8))
* **server:** Add observability system for control execution tracking ([#44](#44)) ([fd0bddc](fd0bddc))
* **server:** add prometheus metrics for endpoints ([#68](#68)) ([775612c](775612c))
* **server:** add time-series stats and split API endpoints ([#6](#6)) ([a0fa597](a0fa597))
* **server:** hard-cut migrate to remove agent UUID ([#44](#44)) ([ee322c9](ee322c9))
* **server:** Optional Policy and many to many relationships ([#41](#41)) ([1a62746](1a62746))
* **ui:** add sql, luna2, json control forms and restructure the code ([#54](#54)) ([c4c1d4a](c4c1d4a))
* **ui:** allow to delete control ([#39](#39)) ([7dc4ca3](7dc4ca3))
* **ui:** Control Store Flow Updated ([#4](#4)) ([dda9f70](dda9f70))
* **ui:** stats dashboard ([#80](#80)) ([4cbb7fe](4cbb7fe))
* **ui:** Steps dropdown rendered based on api return values ([#36](#36)) ([a2aca43](a2aca43))
* **ui:** tests added and some minor ui changes, added error boundaries ([#61](#61)) ([009852b](009852b))
* **ui:** update agent control icon and favicon ([#42](#42)) ([19af8fa](19af8fa))
### Bug Fixes
* **ci:** Add ui scope to PR title validation ([#59](#59)) ([e0fdb52](e0fdb52))
* **ci:** correct galileo contrib path in release build script ([#51](#51)) ([2de6013](2de6013))
* **ci:** Enable pr title on prs ([#56](#56)) ([3d8b5fe](3d8b5fe))
* **ci:** Fix release ([#11](#11)) ([9dd3dd7](9dd3dd7))
* **ci:** Use galileo-automation bot for releases ([#57](#57)) ([bc8eea0](bc8eea0))
* **docs:** Add Example for Evaluator Extension ([#3](#3)) ([c2a70b3](c2a70b3))
* **docs:** add setup script ([#49](#49)) ([7a212c3](7a212c3))
* **docs:** Clean up Protect ([#76](#76)) ([99c16fd](99c16fd))
* **docs:** Fix Examples for LangGraph ([#64](#64)) ([23b30ae](23b30ae))
* **docs:** Improve documentation for open source release ([#47](#47)) ([9018fb3](9018fb3))
* **docs:** Remove old/unused examples ([#66](#66)) ([f417781](f417781))
* **docs:** Update Contributing Guide ([#8](#8)) ([10b34c8](10b34c8))
* **docs:** Update readme ([#37](#37)) ([7531d83](7531d83))
* **docs:** Update README ([#2](#2)) ([379bb15](379bb15))
* **examples:** Control sets cleanup with signed ([#65](#65)) ([af7b5fb](af7b5fb))
* **examples:** Update crew ai example to use evaluator ([#93](#93)) ([1c65084](1c65084))
* **infra:** Add plugins directory to Dockerfile ([#58](#58)) ([171d459](171d459))
* **infra:** install engine/evaluators in server image ([#14](#14)) ([d5ae157](d5ae157))
* **models:** use StrEnum for error enums ([#12](#12)) ([3f41c9f](3f41c9f))
* **sdk-ts:** add conventional commits preset dependency ([#55](#55)) ([540fe9d](540fe9d))
* **sdk-ts:** export npm token for semantic-release npm auth ([#54](#54)) ([1b6b993](1b6b993))
* **sdk:** 54253 add steer action and example ([#38](#38)) ([bf2380a](bf2380a))
* **sdk:** a bug in docker file ([#46](#46)) ([12d1794](12d1794))
* **sdk:** Add step_name as parameter to control ([#25](#25)) ([19ade9d](19ade9d))
* **sdk:** emit observability events for SDK-evaluated controls and fix non_matches propagation ([#24](#24)) ([6a9da69](6a9da69))
* **sdk:** enforce UUID agent IDs ([#9](#9)) ([5ccdbd0](5ccdbd0))
* **sdk:** Fix logging ([#77](#77)) ([b1f078c](b1f078c))
* **sdk:** plugin to evaluator.. agent_protect to agent_control ([#88](#88)) ([fc9b088](fc9b088))
* **server:** enforce public-safe API error responses ([#20](#20)) ([e50d817](e50d817))
* **server:** Feature/56688 fix docker and create bash ([#45](#45)) ([7277e27](7277e27))
* **server:** Feature/56688 fix image bug ([#48](#48)) ([71e6b44](71e6b44))
* **server:** fix alembic migrations ([#47](#47)) ([c19c17c](c19c17c))
* **server:** reject initAgent UUID/name mismatch ([#13](#13)) ([19d61ff](19d61ff))
* tighten evaluation error handling and preserve control data ([52a1ef8](52a1ef8))
* **ui:** Fix UI and clients for simplified step schema ([#75](#75)) ([be2aaf0](be2aaf0))
* **ui:** json validation ([#10](#10)) ([a0cd5af](a0cd5af))
* **ui:** selector subpaths issue ([#34](#34)) ([79cb776](79cb776))
* **ui:** UI feedback fixes ([#27](#27)) ([6004761](6004761))
### Code Refactoring
* **evaluators:** rename plugin to evaluator throughout ([#81](#81)) ([0134682](0134682))
* **evaluators:** split into builtin + extra packages for PyPI ([#5](#5)) ([0e0a78a](0e0a78a))
* **models:** simplify step model and schema ([#70](#70)) ([4c1d637](4c1d637))
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Mar 21, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Jun 12, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@namrataghadi-galileo@lan17@nachiket-galileo@abhinav-galileo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(docs): Add Example for Evaluator Extension - #3

Merged
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example
Feb 3, 2026
Merged

fix(docs): Add Example for Evaluator Extension#3
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example

Conversation

@namrataghadi-galileo

Copy link
Copy Markdown
Contributor

Deep Eval Evaluator example to show how Evaluator can be extended to create custom evaluators.

abhinav-galileoand others added 6 commits January 28, 2026 22:04
- Rename `plugins/` package to `evaluators/` with updated class names:
- RegexPlugin → RegexEvaluator
- ListPlugin → ListEvaluator
- JSONControlEvaluatorPlugin → JSONEvaluator
- SQLControlEvaluatorPlugin → SQLEvaluator
- Luna2Plugin → Luna2Evaluator
- Rename config classes to follow XEvaluatorConfig pattern:
- RegexConfig → RegexEvaluatorConfig
- ListConfig → ListEvaluatorConfig
- JSONControlEvaluatorPluginConfig → JSONEvaluatorConfig
- SQLControlEvaluatorPluginConfig → SQLEvaluatorConfig
- Luna2Config → Luna2EvaluatorConfig
- Update models package:
- Rename plugin.py to evaluator.py
- PluginMetadata → EvaluatorMetadata
- PluginEvaluator → Evaluator
- register_plugin → register_evaluator
- EvaluatorConfig.plugin field → EvaluatorConfig.name
- Update engine package:
- discover_plugins → discover_evaluators
- list_plugins → list_evaluators
- Entry point: agent_control.plugins → agent_control.evaluators
- get_evaluator → get_evaluator_instance (to avoid name collision)
- Update server:
- /api/v1/plugins endpoint → /api/v1/evaluators
- Add migration to rename plugin column to evaluator in evaluator_configs table
- Remove all backwards compatibility aliases
- Update all documentation, examples, and tests
@codecov

codecovBot commented Jan 31, 2026

Copy link
Copy Markdown

The author of this PR, namrataghadi-galileo, is not an activated member of this organization on Codecov.
Please activate this user on Codecov to display this PR comment.
Coverage data is still being uploaded to Codecov.io for purposes of overall coverage calculations.
Please don't hesitate to email us at support@codecov.io with any questions.

Comment threadexamples/deepeval/evaluator.py Outdated
Comment threadexamples/deepeval/pyproject.toml Outdated
@namrataghadi-galileo
namrataghadi-galileo merged commit c2a70b3 into mainFeb 3, 2026
5 checks passed
@namrataghadi-galileo
namrataghadi-galileo deleted the feature/add-deepeval-evaluator-example branch February 3, 2026 22:19
galileo-automation pushed a commit that referenced this pull request Mar 4, 2026
## 1.0.0 (2026-03-04)
### ⚠ BREAKING CHANGES
* **server:** Feature/56688 fix image bug (#48)
* **sdk:** a bug in docker file (#46)
* **server:** Feature/56688 fix docker and create bash (#45)
* **evaluators:** Evaluator reorganization with new package structure
Package Structure:
- agent-control-evaluators (v3.0.0): core + regex, list, json, sql
- agent-control-evaluator-galileo (v3.0.0): Luna2 evaluator
Key Changes:
- Entry points for evaluator discovery (agent_control.evaluators)
- Dot notation for external evaluators (galileo.luna2 not galileo/luna2)
- Dynamic __version__ via importlib.metadata
- Server uses evaluators as runtime dep (no longer vendored)
- Release workflow publishes both packages to PyPI
Bug Fixes:
- JSON evaluator: field_constraints/field_patterns in extra-fields allow-list
- SQL evaluator: LIMIT/OFFSET bypass fix
Migration:
- Import: agent_control_evaluator_galileo.luna2 (not agent_control_evaluators.galileo_luna2)
- DB: UPDATE controls SET evaluator.name replace('/', '.')
* **server:** add time-series stats and split API endpoints (#6)
* **evaluators:** rename plugin to evaluator throughout (#81)
* **models:** simplify step model and schema (#70)
### Features
* Add plugin auto-discovery via Python entry points ([#49](#49)) ([1521182](1521182))
* **docs:** add GitHub badges and CI coverage reporting ([#90](#90)) ([be1fa14](be1fa14))
* **evaluators:** add required_column_values for multi-tenant SQL validation ([#30](#30)) ([532386c](532386c))
* **sdk-ts:** automate semantic-release for npm publishing ([#52](#52)) ([2b43958](2b43958))
* **sdk:** Add PyPI packaging with semantic release ([#52](#52)) ([7c24f7f](7c24f7f))
* **sdk:** Auto-populate init() steps from [@control](https://github.com/control)() decorators ([#23](#23)) ([dc0f2a4](dc0f2a4))
* **sdk:** export ControlScope, ControlMatch, and EvaluatorResult models ([#18](#18)) ([0d49cad](0d49cad))
* **sdk:** Get Agent Controls from SDK Init ([#15](#15)) ([a485f93](a485f93))
* **sdk:** Refresh controls in a background loop ([#43](#43)) ([03f826d](03f826d))
* **sdk:** ship TypeScript SDK with deterministic method naming ([#32](#32)) ([a76e9b0](a76e9b0))
* **server:** add evaluator config store ([#78](#78)) ([cc14aa6](cc14aa6))
* **server:** add initAgent conflict_mode overwrite mode with SDK defaults ([#40](#40)) ([f3ed2b8](f3ed2b8))
* **server:** Add observability system for control execution tracking ([#44](#44)) ([fd0bddc](fd0bddc))
* **server:** add prometheus metrics for endpoints ([#68](#68)) ([775612c](775612c))
* **server:** add time-series stats and split API endpoints ([#6](#6)) ([a0fa597](a0fa597))
* **server:** hard-cut migrate to remove agent UUID ([#44](#44)) ([ee322c9](ee322c9))
* **server:** Optional Policy and many to many relationships ([#41](#41)) ([1a62746](1a62746))
* **ui:** add sql, luna2, json control forms and restructure the code ([#54](#54)) ([c4c1d4a](c4c1d4a))
* **ui:** allow to delete control ([#39](#39)) ([7dc4ca3](7dc4ca3))
* **ui:** Control Store Flow Updated ([#4](#4)) ([dda9f70](dda9f70))
* **ui:** stats dashboard ([#80](#80)) ([4cbb7fe](4cbb7fe))
* **ui:** Steps dropdown rendered based on api return values ([#36](#36)) ([a2aca43](a2aca43))
* **ui:** tests added and some minor ui changes, added error boundaries ([#61](#61)) ([009852b](009852b))
* **ui:** update agent control icon and favicon ([#42](#42)) ([19af8fa](19af8fa))
### Bug Fixes
* **ci:** Add ui scope to PR title validation ([#59](#59)) ([e0fdb52](e0fdb52))
* **ci:** correct galileo contrib path in release build script ([#51](#51)) ([2de6013](2de6013))
* **ci:** Enable pr title on prs ([#56](#56)) ([3d8b5fe](3d8b5fe))
* **ci:** Fix release ([#11](#11)) ([9dd3dd7](9dd3dd7))
* **ci:** Use galileo-automation bot for releases ([#57](#57)) ([bc8eea0](bc8eea0))
* **docs:** Add Example for Evaluator Extension ([#3](#3)) ([c2a70b3](c2a70b3))
* **docs:** add setup script ([#49](#49)) ([7a212c3](7a212c3))
* **docs:** Clean up Protect ([#76](#76)) ([99c16fd](99c16fd))
* **docs:** Fix Examples for LangGraph ([#64](#64)) ([23b30ae](23b30ae))
* **docs:** Improve documentation for open source release ([#47](#47)) ([9018fb3](9018fb3))
* **docs:** Remove old/unused examples ([#66](#66)) ([f417781](f417781))
* **docs:** Update Contributing Guide ([#8](#8)) ([10b34c8](10b34c8))
* **docs:** Update readme ([#37](#37)) ([7531d83](7531d83))
* **docs:** Update README ([#2](#2)) ([379bb15](379bb15))
* **examples:** Control sets cleanup with signed ([#65](#65)) ([af7b5fb](af7b5fb))
* **examples:** Update crew ai example to use evaluator ([#93](#93)) ([1c65084](1c65084))
* **infra:** Add plugins directory to Dockerfile ([#58](#58)) ([171d459](171d459))
* **infra:** install engine/evaluators in server image ([#14](#14)) ([d5ae157](d5ae157))
* **models:** use StrEnum for error enums ([#12](#12)) ([3f41c9f](3f41c9f))
* **sdk-ts:** add conventional commits preset dependency ([#55](#55)) ([540fe9d](540fe9d))
* **sdk-ts:** export npm token for semantic-release npm auth ([#54](#54)) ([1b6b993](1b6b993))
* **sdk:** 54253 add steer action and example ([#38](#38)) ([bf2380a](bf2380a))
* **sdk:** a bug in docker file ([#46](#46)) ([12d1794](12d1794))
* **sdk:** Add step_name as parameter to control ([#25](#25)) ([19ade9d](19ade9d))
* **sdk:** emit observability events for SDK-evaluated controls and fix non_matches propagation ([#24](#24)) ([6a9da69](6a9da69))
* **sdk:** enforce UUID agent IDs ([#9](#9)) ([5ccdbd0](5ccdbd0))
* **sdk:** Fix logging ([#77](#77)) ([b1f078c](b1f078c))
* **sdk:** plugin to evaluator.. agent_protect to agent_control ([#88](#88)) ([fc9b088](fc9b088))
* **server:** enforce public-safe API error responses ([#20](#20)) ([e50d817](e50d817))
* **server:** Feature/56688 fix docker and create bash ([#45](#45)) ([7277e27](7277e27))
* **server:** Feature/56688 fix image bug ([#48](#48)) ([71e6b44](71e6b44))
* **server:** fix alembic migrations ([#47](#47)) ([c19c17c](c19c17c))
* **server:** reject initAgent UUID/name mismatch ([#13](#13)) ([19d61ff](19d61ff))
* tighten evaluation error handling and preserve control data ([52a1ef8](52a1ef8))
* **ui:** Fix UI and clients for simplified step schema ([#75](#75)) ([be2aaf0](be2aaf0))
* **ui:** json validation ([#10](#10)) ([a0cd5af](a0cd5af))
* **ui:** selector subpaths issue ([#34](#34)) ([79cb776](79cb776))
* **ui:** UI feedback fixes ([#27](#27)) ([6004761](6004761))
### Code Refactoring
* **evaluators:** rename plugin to evaluator throughout ([#81](#81)) ([0134682](0134682))
* **evaluators:** split into builtin + extra packages for PyPI ([#5](#5)) ([0e0a78a](0e0a78a))
* **models:** simplify step model and schema ([#70](#70)) ([4c1d637](4c1d637))
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Mar 21, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Jun 12, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@namrataghadi-galileo@lan17@nachiket-galileo@abhinav-galileo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(docs): Add Example for Evaluator Extension - #3

Merged
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example
Feb 3, 2026
Merged

fix(docs): Add Example for Evaluator Extension#3
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example

Conversation

@namrataghadi-galileo

Copy link
Copy Markdown
Contributor

Deep Eval Evaluator example to show how Evaluator can be extended to create custom evaluators.

abhinav-galileoand others added 6 commits January 28, 2026 22:04
- Rename `plugins/` package to `evaluators/` with updated class names:
- RegexPlugin → RegexEvaluator
- ListPlugin → ListEvaluator
- JSONControlEvaluatorPlugin → JSONEvaluator
- SQLControlEvaluatorPlugin → SQLEvaluator
- Luna2Plugin → Luna2Evaluator
- Rename config classes to follow XEvaluatorConfig pattern:
- RegexConfig → RegexEvaluatorConfig
- ListConfig → ListEvaluatorConfig
- JSONControlEvaluatorPluginConfig → JSONEvaluatorConfig
- SQLControlEvaluatorPluginConfig → SQLEvaluatorConfig
- Luna2Config → Luna2EvaluatorConfig
- Update models package:
- Rename plugin.py to evaluator.py
- PluginMetadata → EvaluatorMetadata
- PluginEvaluator → Evaluator
- register_plugin → register_evaluator
- EvaluatorConfig.plugin field → EvaluatorConfig.name
- Update engine package:
- discover_plugins → discover_evaluators
- list_plugins → list_evaluators
- Entry point: agent_control.plugins → agent_control.evaluators
- get_evaluator → get_evaluator_instance (to avoid name collision)
- Update server:
- /api/v1/plugins endpoint → /api/v1/evaluators
- Add migration to rename plugin column to evaluator in evaluator_configs table
- Remove all backwards compatibility aliases
- Update all documentation, examples, and tests
@codecov

codecovBot commented Jan 31, 2026

Copy link
Copy Markdown

The author of this PR, namrataghadi-galileo, is not an activated member of this organization on Codecov.
Please activate this user on Codecov to display this PR comment.
Coverage data is still being uploaded to Codecov.io for purposes of overall coverage calculations.
Please don't hesitate to email us at support@codecov.io with any questions.

Comment threadexamples/deepeval/evaluator.py Outdated
Comment threadexamples/deepeval/pyproject.toml Outdated
@namrataghadi-galileo
namrataghadi-galileo merged commit c2a70b3 into mainFeb 3, 2026
5 checks passed
@namrataghadi-galileo
namrataghadi-galileo deleted the feature/add-deepeval-evaluator-example branch February 3, 2026 22:19
galileo-automation pushed a commit that referenced this pull request Mar 4, 2026
## 1.0.0 (2026-03-04)
### ⚠ BREAKING CHANGES
* **server:** Feature/56688 fix image bug (#48)
* **sdk:** a bug in docker file (#46)
* **server:** Feature/56688 fix docker and create bash (#45)
* **evaluators:** Evaluator reorganization with new package structure
Package Structure:
- agent-control-evaluators (v3.0.0): core + regex, list, json, sql
- agent-control-evaluator-galileo (v3.0.0): Luna2 evaluator
Key Changes:
- Entry points for evaluator discovery (agent_control.evaluators)
- Dot notation for external evaluators (galileo.luna2 not galileo/luna2)
- Dynamic __version__ via importlib.metadata
- Server uses evaluators as runtime dep (no longer vendored)
- Release workflow publishes both packages to PyPI
Bug Fixes:
- JSON evaluator: field_constraints/field_patterns in extra-fields allow-list
- SQL evaluator: LIMIT/OFFSET bypass fix
Migration:
- Import: agent_control_evaluator_galileo.luna2 (not agent_control_evaluators.galileo_luna2)
- DB: UPDATE controls SET evaluator.name replace('/', '.')
* **server:** add time-series stats and split API endpoints (#6)
* **evaluators:** rename plugin to evaluator throughout (#81)
* **models:** simplify step model and schema (#70)
### Features
* Add plugin auto-discovery via Python entry points ([#49](#49)) ([1521182](1521182))
* **docs:** add GitHub badges and CI coverage reporting ([#90](#90)) ([be1fa14](be1fa14))
* **evaluators:** add required_column_values for multi-tenant SQL validation ([#30](#30)) ([532386c](532386c))
* **sdk-ts:** automate semantic-release for npm publishing ([#52](#52)) ([2b43958](2b43958))
* **sdk:** Add PyPI packaging with semantic release ([#52](#52)) ([7c24f7f](7c24f7f))
* **sdk:** Auto-populate init() steps from [@control](https://github.com/control)() decorators ([#23](#23)) ([dc0f2a4](dc0f2a4))
* **sdk:** export ControlScope, ControlMatch, and EvaluatorResult models ([#18](#18)) ([0d49cad](0d49cad))
* **sdk:** Get Agent Controls from SDK Init ([#15](#15)) ([a485f93](a485f93))
* **sdk:** Refresh controls in a background loop ([#43](#43)) ([03f826d](03f826d))
* **sdk:** ship TypeScript SDK with deterministic method naming ([#32](#32)) ([a76e9b0](a76e9b0))
* **server:** add evaluator config store ([#78](#78)) ([cc14aa6](cc14aa6))
* **server:** add initAgent conflict_mode overwrite mode with SDK defaults ([#40](#40)) ([f3ed2b8](f3ed2b8))
* **server:** Add observability system for control execution tracking ([#44](#44)) ([fd0bddc](fd0bddc))
* **server:** add prometheus metrics for endpoints ([#68](#68)) ([775612c](775612c))
* **server:** add time-series stats and split API endpoints ([#6](#6)) ([a0fa597](a0fa597))
* **server:** hard-cut migrate to remove agent UUID ([#44](#44)) ([ee322c9](ee322c9))
* **server:** Optional Policy and many to many relationships ([#41](#41)) ([1a62746](1a62746))
* **ui:** add sql, luna2, json control forms and restructure the code ([#54](#54)) ([c4c1d4a](c4c1d4a))
* **ui:** allow to delete control ([#39](#39)) ([7dc4ca3](7dc4ca3))
* **ui:** Control Store Flow Updated ([#4](#4)) ([dda9f70](dda9f70))
* **ui:** stats dashboard ([#80](#80)) ([4cbb7fe](4cbb7fe))
* **ui:** Steps dropdown rendered based on api return values ([#36](#36)) ([a2aca43](a2aca43))
* **ui:** tests added and some minor ui changes, added error boundaries ([#61](#61)) ([009852b](009852b))
* **ui:** update agent control icon and favicon ([#42](#42)) ([19af8fa](19af8fa))
### Bug Fixes
* **ci:** Add ui scope to PR title validation ([#59](#59)) ([e0fdb52](e0fdb52))
* **ci:** correct galileo contrib path in release build script ([#51](#51)) ([2de6013](2de6013))
* **ci:** Enable pr title on prs ([#56](#56)) ([3d8b5fe](3d8b5fe))
* **ci:** Fix release ([#11](#11)) ([9dd3dd7](9dd3dd7))
* **ci:** Use galileo-automation bot for releases ([#57](#57)) ([bc8eea0](bc8eea0))
* **docs:** Add Example for Evaluator Extension ([#3](#3)) ([c2a70b3](c2a70b3))
* **docs:** add setup script ([#49](#49)) ([7a212c3](7a212c3))
* **docs:** Clean up Protect ([#76](#76)) ([99c16fd](99c16fd))
* **docs:** Fix Examples for LangGraph ([#64](#64)) ([23b30ae](23b30ae))
* **docs:** Improve documentation for open source release ([#47](#47)) ([9018fb3](9018fb3))
* **docs:** Remove old/unused examples ([#66](#66)) ([f417781](f417781))
* **docs:** Update Contributing Guide ([#8](#8)) ([10b34c8](10b34c8))
* **docs:** Update readme ([#37](#37)) ([7531d83](7531d83))
* **docs:** Update README ([#2](#2)) ([379bb15](379bb15))
* **examples:** Control sets cleanup with signed ([#65](#65)) ([af7b5fb](af7b5fb))
* **examples:** Update crew ai example to use evaluator ([#93](#93)) ([1c65084](1c65084))
* **infra:** Add plugins directory to Dockerfile ([#58](#58)) ([171d459](171d459))
* **infra:** install engine/evaluators in server image ([#14](#14)) ([d5ae157](d5ae157))
* **models:** use StrEnum for error enums ([#12](#12)) ([3f41c9f](3f41c9f))
* **sdk-ts:** add conventional commits preset dependency ([#55](#55)) ([540fe9d](540fe9d))
* **sdk-ts:** export npm token for semantic-release npm auth ([#54](#54)) ([1b6b993](1b6b993))
* **sdk:** 54253 add steer action and example ([#38](#38)) ([bf2380a](bf2380a))
* **sdk:** a bug in docker file ([#46](#46)) ([12d1794](12d1794))
* **sdk:** Add step_name as parameter to control ([#25](#25)) ([19ade9d](19ade9d))
* **sdk:** emit observability events for SDK-evaluated controls and fix non_matches propagation ([#24](#24)) ([6a9da69](6a9da69))
* **sdk:** enforce UUID agent IDs ([#9](#9)) ([5ccdbd0](5ccdbd0))
* **sdk:** Fix logging ([#77](#77)) ([b1f078c](b1f078c))
* **sdk:** plugin to evaluator.. agent_protect to agent_control ([#88](#88)) ([fc9b088](fc9b088))
* **server:** enforce public-safe API error responses ([#20](#20)) ([e50d817](e50d817))
* **server:** Feature/56688 fix docker and create bash ([#45](#45)) ([7277e27](7277e27))
* **server:** Feature/56688 fix image bug ([#48](#48)) ([71e6b44](71e6b44))
* **server:** fix alembic migrations ([#47](#47)) ([c19c17c](c19c17c))
* **server:** reject initAgent UUID/name mismatch ([#13](#13)) ([19d61ff](19d61ff))
* tighten evaluation error handling and preserve control data ([52a1ef8](52a1ef8))
* **ui:** Fix UI and clients for simplified step schema ([#75](#75)) ([be2aaf0](be2aaf0))
* **ui:** json validation ([#10](#10)) ([a0cd5af](a0cd5af))
* **ui:** selector subpaths issue ([#34](#34)) ([79cb776](79cb776))
* **ui:** UI feedback fixes ([#27](#27)) ([6004761](6004761))
### Code Refactoring
* **evaluators:** rename plugin to evaluator throughout ([#81](#81)) ([0134682](0134682))
* **evaluators:** split into builtin + extra packages for PyPI ([#5](#5)) ([0e0a78a](0e0a78a))
* **models:** simplify step model and schema ([#70](#70)) ([4c1d637](4c1d637))
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Mar 21, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Jun 12, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@namrataghadi-galileo@lan17@nachiket-galileo@abhinav-galileo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(docs): Add Example for Evaluator Extension - #3

Merged
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example
Feb 3, 2026
Merged

fix(docs): Add Example for Evaluator Extension#3
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example

Conversation

@namrataghadi-galileo

Copy link
Copy Markdown
Contributor

Deep Eval Evaluator example to show how Evaluator can be extended to create custom evaluators.

abhinav-galileoand others added 6 commits January 28, 2026 22:04
- Rename `plugins/` package to `evaluators/` with updated class names:
- RegexPlugin → RegexEvaluator
- ListPlugin → ListEvaluator
- JSONControlEvaluatorPlugin → JSONEvaluator
- SQLControlEvaluatorPlugin → SQLEvaluator
- Luna2Plugin → Luna2Evaluator
- Rename config classes to follow XEvaluatorConfig pattern:
- RegexConfig → RegexEvaluatorConfig
- ListConfig → ListEvaluatorConfig
- JSONControlEvaluatorPluginConfig → JSONEvaluatorConfig
- SQLControlEvaluatorPluginConfig → SQLEvaluatorConfig
- Luna2Config → Luna2EvaluatorConfig
- Update models package:
- Rename plugin.py to evaluator.py
- PluginMetadata → EvaluatorMetadata
- PluginEvaluator → Evaluator
- register_plugin → register_evaluator
- EvaluatorConfig.plugin field → EvaluatorConfig.name
- Update engine package:
- discover_plugins → discover_evaluators
- list_plugins → list_evaluators
- Entry point: agent_control.plugins → agent_control.evaluators
- get_evaluator → get_evaluator_instance (to avoid name collision)
- Update server:
- /api/v1/plugins endpoint → /api/v1/evaluators
- Add migration to rename plugin column to evaluator in evaluator_configs table
- Remove all backwards compatibility aliases
- Update all documentation, examples, and tests
@codecov

codecovBot commented Jan 31, 2026

Copy link
Copy Markdown

The author of this PR, namrataghadi-galileo, is not an activated member of this organization on Codecov.
Please activate this user on Codecov to display this PR comment.
Coverage data is still being uploaded to Codecov.io for purposes of overall coverage calculations.
Please don't hesitate to email us at support@codecov.io with any questions.

Comment threadexamples/deepeval/evaluator.py Outdated
Comment threadexamples/deepeval/pyproject.toml Outdated
@namrataghadi-galileo
namrataghadi-galileo merged commit c2a70b3 into mainFeb 3, 2026
5 checks passed
@namrataghadi-galileo
namrataghadi-galileo deleted the feature/add-deepeval-evaluator-example branch February 3, 2026 22:19
galileo-automation pushed a commit that referenced this pull request Mar 4, 2026
## 1.0.0 (2026-03-04)
### ⚠ BREAKING CHANGES
* **server:** Feature/56688 fix image bug (#48)
* **sdk:** a bug in docker file (#46)
* **server:** Feature/56688 fix docker and create bash (#45)
* **evaluators:** Evaluator reorganization with new package structure
Package Structure:
- agent-control-evaluators (v3.0.0): core + regex, list, json, sql
- agent-control-evaluator-galileo (v3.0.0): Luna2 evaluator
Key Changes:
- Entry points for evaluator discovery (agent_control.evaluators)
- Dot notation for external evaluators (galileo.luna2 not galileo/luna2)
- Dynamic __version__ via importlib.metadata
- Server uses evaluators as runtime dep (no longer vendored)
- Release workflow publishes both packages to PyPI
Bug Fixes:
- JSON evaluator: field_constraints/field_patterns in extra-fields allow-list
- SQL evaluator: LIMIT/OFFSET bypass fix
Migration:
- Import: agent_control_evaluator_galileo.luna2 (not agent_control_evaluators.galileo_luna2)
- DB: UPDATE controls SET evaluator.name replace('/', '.')
* **server:** add time-series stats and split API endpoints (#6)
* **evaluators:** rename plugin to evaluator throughout (#81)
* **models:** simplify step model and schema (#70)
### Features
* Add plugin auto-discovery via Python entry points ([#49](#49)) ([1521182](1521182))
* **docs:** add GitHub badges and CI coverage reporting ([#90](#90)) ([be1fa14](be1fa14))
* **evaluators:** add required_column_values for multi-tenant SQL validation ([#30](#30)) ([532386c](532386c))
* **sdk-ts:** automate semantic-release for npm publishing ([#52](#52)) ([2b43958](2b43958))
* **sdk:** Add PyPI packaging with semantic release ([#52](#52)) ([7c24f7f](7c24f7f))
* **sdk:** Auto-populate init() steps from [@control](https://github.com/control)() decorators ([#23](#23)) ([dc0f2a4](dc0f2a4))
* **sdk:** export ControlScope, ControlMatch, and EvaluatorResult models ([#18](#18)) ([0d49cad](0d49cad))
* **sdk:** Get Agent Controls from SDK Init ([#15](#15)) ([a485f93](a485f93))
* **sdk:** Refresh controls in a background loop ([#43](#43)) ([03f826d](03f826d))
* **sdk:** ship TypeScript SDK with deterministic method naming ([#32](#32)) ([a76e9b0](a76e9b0))
* **server:** add evaluator config store ([#78](#78)) ([cc14aa6](cc14aa6))
* **server:** add initAgent conflict_mode overwrite mode with SDK defaults ([#40](#40)) ([f3ed2b8](f3ed2b8))
* **server:** Add observability system for control execution tracking ([#44](#44)) ([fd0bddc](fd0bddc))
* **server:** add prometheus metrics for endpoints ([#68](#68)) ([775612c](775612c))
* **server:** add time-series stats and split API endpoints ([#6](#6)) ([a0fa597](a0fa597))
* **server:** hard-cut migrate to remove agent UUID ([#44](#44)) ([ee322c9](ee322c9))
* **server:** Optional Policy and many to many relationships ([#41](#41)) ([1a62746](1a62746))
* **ui:** add sql, luna2, json control forms and restructure the code ([#54](#54)) ([c4c1d4a](c4c1d4a))
* **ui:** allow to delete control ([#39](#39)) ([7dc4ca3](7dc4ca3))
* **ui:** Control Store Flow Updated ([#4](#4)) ([dda9f70](dda9f70))
* **ui:** stats dashboard ([#80](#80)) ([4cbb7fe](4cbb7fe))
* **ui:** Steps dropdown rendered based on api return values ([#36](#36)) ([a2aca43](a2aca43))
* **ui:** tests added and some minor ui changes, added error boundaries ([#61](#61)) ([009852b](009852b))
* **ui:** update agent control icon and favicon ([#42](#42)) ([19af8fa](19af8fa))
### Bug Fixes
* **ci:** Add ui scope to PR title validation ([#59](#59)) ([e0fdb52](e0fdb52))
* **ci:** correct galileo contrib path in release build script ([#51](#51)) ([2de6013](2de6013))
* **ci:** Enable pr title on prs ([#56](#56)) ([3d8b5fe](3d8b5fe))
* **ci:** Fix release ([#11](#11)) ([9dd3dd7](9dd3dd7))
* **ci:** Use galileo-automation bot for releases ([#57](#57)) ([bc8eea0](bc8eea0))
* **docs:** Add Example for Evaluator Extension ([#3](#3)) ([c2a70b3](c2a70b3))
* **docs:** add setup script ([#49](#49)) ([7a212c3](7a212c3))
* **docs:** Clean up Protect ([#76](#76)) ([99c16fd](99c16fd))
* **docs:** Fix Examples for LangGraph ([#64](#64)) ([23b30ae](23b30ae))
* **docs:** Improve documentation for open source release ([#47](#47)) ([9018fb3](9018fb3))
* **docs:** Remove old/unused examples ([#66](#66)) ([f417781](f417781))
* **docs:** Update Contributing Guide ([#8](#8)) ([10b34c8](10b34c8))
* **docs:** Update readme ([#37](#37)) ([7531d83](7531d83))
* **docs:** Update README ([#2](#2)) ([379bb15](379bb15))
* **examples:** Control sets cleanup with signed ([#65](#65)) ([af7b5fb](af7b5fb))
* **examples:** Update crew ai example to use evaluator ([#93](#93)) ([1c65084](1c65084))
* **infra:** Add plugins directory to Dockerfile ([#58](#58)) ([171d459](171d459))
* **infra:** install engine/evaluators in server image ([#14](#14)) ([d5ae157](d5ae157))
* **models:** use StrEnum for error enums ([#12](#12)) ([3f41c9f](3f41c9f))
* **sdk-ts:** add conventional commits preset dependency ([#55](#55)) ([540fe9d](540fe9d))
* **sdk-ts:** export npm token for semantic-release npm auth ([#54](#54)) ([1b6b993](1b6b993))
* **sdk:** 54253 add steer action and example ([#38](#38)) ([bf2380a](bf2380a))
* **sdk:** a bug in docker file ([#46](#46)) ([12d1794](12d1794))
* **sdk:** Add step_name as parameter to control ([#25](#25)) ([19ade9d](19ade9d))
* **sdk:** emit observability events for SDK-evaluated controls and fix non_matches propagation ([#24](#24)) ([6a9da69](6a9da69))
* **sdk:** enforce UUID agent IDs ([#9](#9)) ([5ccdbd0](5ccdbd0))
* **sdk:** Fix logging ([#77](#77)) ([b1f078c](b1f078c))
* **sdk:** plugin to evaluator.. agent_protect to agent_control ([#88](#88)) ([fc9b088](fc9b088))
* **server:** enforce public-safe API error responses ([#20](#20)) ([e50d817](e50d817))
* **server:** Feature/56688 fix docker and create bash ([#45](#45)) ([7277e27](7277e27))
* **server:** Feature/56688 fix image bug ([#48](#48)) ([71e6b44](71e6b44))
* **server:** fix alembic migrations ([#47](#47)) ([c19c17c](c19c17c))
* **server:** reject initAgent UUID/name mismatch ([#13](#13)) ([19d61ff](19d61ff))
* tighten evaluation error handling and preserve control data ([52a1ef8](52a1ef8))
* **ui:** Fix UI and clients for simplified step schema ([#75](#75)) ([be2aaf0](be2aaf0))
* **ui:** json validation ([#10](#10)) ([a0cd5af](a0cd5af))
* **ui:** selector subpaths issue ([#34](#34)) ([79cb776](79cb776))
* **ui:** UI feedback fixes ([#27](#27)) ([6004761](6004761))
### Code Refactoring
* **evaluators:** rename plugin to evaluator throughout ([#81](#81)) ([0134682](0134682))
* **evaluators:** split into builtin + extra packages for PyPI ([#5](#5)) ([0e0a78a](0e0a78a))
* **models:** simplify step model and schema ([#70](#70)) ([4c1d637](4c1d637))
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Mar 21, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Jun 12, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@namrataghadi-galileo@lan17@nachiket-galileo@abhinav-galileo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(docs): Add Example for Evaluator Extension - #3

Merged
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example
Feb 3, 2026
Merged

fix(docs): Add Example for Evaluator Extension#3
namrataghadi-galileo merged 11 commits into
mainfrom
feature/add-deepeval-evaluator-example

Conversation

@namrataghadi-galileo

Copy link
Copy Markdown
Contributor

Deep Eval Evaluator example to show how Evaluator can be extended to create custom evaluators.

abhinav-galileoand others added 6 commits January 28, 2026 22:04
- Rename `plugins/` package to `evaluators/` with updated class names:
- RegexPlugin → RegexEvaluator
- ListPlugin → ListEvaluator
- JSONControlEvaluatorPlugin → JSONEvaluator
- SQLControlEvaluatorPlugin → SQLEvaluator
- Luna2Plugin → Luna2Evaluator
- Rename config classes to follow XEvaluatorConfig pattern:
- RegexConfig → RegexEvaluatorConfig
- ListConfig → ListEvaluatorConfig
- JSONControlEvaluatorPluginConfig → JSONEvaluatorConfig
- SQLControlEvaluatorPluginConfig → SQLEvaluatorConfig
- Luna2Config → Luna2EvaluatorConfig
- Update models package:
- Rename plugin.py to evaluator.py
- PluginMetadata → EvaluatorMetadata
- PluginEvaluator → Evaluator
- register_plugin → register_evaluator
- EvaluatorConfig.plugin field → EvaluatorConfig.name
- Update engine package:
- discover_plugins → discover_evaluators
- list_plugins → list_evaluators
- Entry point: agent_control.plugins → agent_control.evaluators
- get_evaluator → get_evaluator_instance (to avoid name collision)
- Update server:
- /api/v1/plugins endpoint → /api/v1/evaluators
- Add migration to rename plugin column to evaluator in evaluator_configs table
- Remove all backwards compatibility aliases
- Update all documentation, examples, and tests
@codecov

codecovBot commented Jan 31, 2026

Copy link
Copy Markdown

The author of this PR, namrataghadi-galileo, is not an activated member of this organization on Codecov.
Please activate this user on Codecov to display this PR comment.
Coverage data is still being uploaded to Codecov.io for purposes of overall coverage calculations.
Please don't hesitate to email us at support@codecov.io with any questions.

Comment threadexamples/deepeval/evaluator.py Outdated
Comment threadexamples/deepeval/pyproject.toml Outdated
@namrataghadi-galileo
namrataghadi-galileo merged commit c2a70b3 into mainFeb 3, 2026
5 checks passed
@namrataghadi-galileo
namrataghadi-galileo deleted the feature/add-deepeval-evaluator-example branch February 3, 2026 22:19
galileo-automation pushed a commit that referenced this pull request Mar 4, 2026
## 1.0.0 (2026-03-04)
### ⚠ BREAKING CHANGES
* **server:** Feature/56688 fix image bug (#48)
* **sdk:** a bug in docker file (#46)
* **server:** Feature/56688 fix docker and create bash (#45)
* **evaluators:** Evaluator reorganization with new package structure
Package Structure:
- agent-control-evaluators (v3.0.0): core + regex, list, json, sql
- agent-control-evaluator-galileo (v3.0.0): Luna2 evaluator
Key Changes:
- Entry points for evaluator discovery (agent_control.evaluators)
- Dot notation for external evaluators (galileo.luna2 not galileo/luna2)
- Dynamic __version__ via importlib.metadata
- Server uses evaluators as runtime dep (no longer vendored)
- Release workflow publishes both packages to PyPI
Bug Fixes:
- JSON evaluator: field_constraints/field_patterns in extra-fields allow-list
- SQL evaluator: LIMIT/OFFSET bypass fix
Migration:
- Import: agent_control_evaluator_galileo.luna2 (not agent_control_evaluators.galileo_luna2)
- DB: UPDATE controls SET evaluator.name replace('/', '.')
* **server:** add time-series stats and split API endpoints (#6)
* **evaluators:** rename plugin to evaluator throughout (#81)
* **models:** simplify step model and schema (#70)
### Features
* Add plugin auto-discovery via Python entry points ([#49](#49)) ([1521182](1521182))
* **docs:** add GitHub badges and CI coverage reporting ([#90](#90)) ([be1fa14](be1fa14))
* **evaluators:** add required_column_values for multi-tenant SQL validation ([#30](#30)) ([532386c](532386c))
* **sdk-ts:** automate semantic-release for npm publishing ([#52](#52)) ([2b43958](2b43958))
* **sdk:** Add PyPI packaging with semantic release ([#52](#52)) ([7c24f7f](7c24f7f))
* **sdk:** Auto-populate init() steps from [@control](https://github.com/control)() decorators ([#23](#23)) ([dc0f2a4](dc0f2a4))
* **sdk:** export ControlScope, ControlMatch, and EvaluatorResult models ([#18](#18)) ([0d49cad](0d49cad))
* **sdk:** Get Agent Controls from SDK Init ([#15](#15)) ([a485f93](a485f93))
* **sdk:** Refresh controls in a background loop ([#43](#43)) ([03f826d](03f826d))
* **sdk:** ship TypeScript SDK with deterministic method naming ([#32](#32)) ([a76e9b0](a76e9b0))
* **server:** add evaluator config store ([#78](#78)) ([cc14aa6](cc14aa6))
* **server:** add initAgent conflict_mode overwrite mode with SDK defaults ([#40](#40)) ([f3ed2b8](f3ed2b8))
* **server:** Add observability system for control execution tracking ([#44](#44)) ([fd0bddc](fd0bddc))
* **server:** add prometheus metrics for endpoints ([#68](#68)) ([775612c](775612c))
* **server:** add time-series stats and split API endpoints ([#6](#6)) ([a0fa597](a0fa597))
* **server:** hard-cut migrate to remove agent UUID ([#44](#44)) ([ee322c9](ee322c9))
* **server:** Optional Policy and many to many relationships ([#41](#41)) ([1a62746](1a62746))
* **ui:** add sql, luna2, json control forms and restructure the code ([#54](#54)) ([c4c1d4a](c4c1d4a))
* **ui:** allow to delete control ([#39](#39)) ([7dc4ca3](7dc4ca3))
* **ui:** Control Store Flow Updated ([#4](#4)) ([dda9f70](dda9f70))
* **ui:** stats dashboard ([#80](#80)) ([4cbb7fe](4cbb7fe))
* **ui:** Steps dropdown rendered based on api return values ([#36](#36)) ([a2aca43](a2aca43))
* **ui:** tests added and some minor ui changes, added error boundaries ([#61](#61)) ([009852b](009852b))
* **ui:** update agent control icon and favicon ([#42](#42)) ([19af8fa](19af8fa))
### Bug Fixes
* **ci:** Add ui scope to PR title validation ([#59](#59)) ([e0fdb52](e0fdb52))
* **ci:** correct galileo contrib path in release build script ([#51](#51)) ([2de6013](2de6013))
* **ci:** Enable pr title on prs ([#56](#56)) ([3d8b5fe](3d8b5fe))
* **ci:** Fix release ([#11](#11)) ([9dd3dd7](9dd3dd7))
* **ci:** Use galileo-automation bot for releases ([#57](#57)) ([bc8eea0](bc8eea0))
* **docs:** Add Example for Evaluator Extension ([#3](#3)) ([c2a70b3](c2a70b3))
* **docs:** add setup script ([#49](#49)) ([7a212c3](7a212c3))
* **docs:** Clean up Protect ([#76](#76)) ([99c16fd](99c16fd))
* **docs:** Fix Examples for LangGraph ([#64](#64)) ([23b30ae](23b30ae))
* **docs:** Improve documentation for open source release ([#47](#47)) ([9018fb3](9018fb3))
* **docs:** Remove old/unused examples ([#66](#66)) ([f417781](f417781))
* **docs:** Update Contributing Guide ([#8](#8)) ([10b34c8](10b34c8))
* **docs:** Update readme ([#37](#37)) ([7531d83](7531d83))
* **docs:** Update README ([#2](#2)) ([379bb15](379bb15))
* **examples:** Control sets cleanup with signed ([#65](#65)) ([af7b5fb](af7b5fb))
* **examples:** Update crew ai example to use evaluator ([#93](#93)) ([1c65084](1c65084))
* **infra:** Add plugins directory to Dockerfile ([#58](#58)) ([171d459](171d459))
* **infra:** install engine/evaluators in server image ([#14](#14)) ([d5ae157](d5ae157))
* **models:** use StrEnum for error enums ([#12](#12)) ([3f41c9f](3f41c9f))
* **sdk-ts:** add conventional commits preset dependency ([#55](#55)) ([540fe9d](540fe9d))
* **sdk-ts:** export npm token for semantic-release npm auth ([#54](#54)) ([1b6b993](1b6b993))
* **sdk:** 54253 add steer action and example ([#38](#38)) ([bf2380a](bf2380a))
* **sdk:** a bug in docker file ([#46](#46)) ([12d1794](12d1794))
* **sdk:** Add step_name as parameter to control ([#25](#25)) ([19ade9d](19ade9d))
* **sdk:** emit observability events for SDK-evaluated controls and fix non_matches propagation ([#24](#24)) ([6a9da69](6a9da69))
* **sdk:** enforce UUID agent IDs ([#9](#9)) ([5ccdbd0](5ccdbd0))
* **sdk:** Fix logging ([#77](#77)) ([b1f078c](b1f078c))
* **sdk:** plugin to evaluator.. agent_protect to agent_control ([#88](#88)) ([fc9b088](fc9b088))
* **server:** enforce public-safe API error responses ([#20](#20)) ([e50d817](e50d817))
* **server:** Feature/56688 fix docker and create bash ([#45](#45)) ([7277e27](7277e27))
* **server:** Feature/56688 fix image bug ([#48](#48)) ([71e6b44](71e6b44))
* **server:** fix alembic migrations ([#47](#47)) ([c19c17c](c19c17c))
* **server:** reject initAgent UUID/name mismatch ([#13](#13)) ([19d61ff](19d61ff))
* tighten evaluation error handling and preserve control data ([52a1ef8](52a1ef8))
* **ui:** Fix UI and clients for simplified step schema ([#75](#75)) ([be2aaf0](be2aaf0))
* **ui:** json validation ([#10](#10)) ([a0cd5af](a0cd5af))
* **ui:** selector subpaths issue ([#34](#34)) ([79cb776](79cb776))
* **ui:** UI feedback fixes ([#27](#27)) ([6004761](6004761))
### Code Refactoring
* **evaluators:** rename plugin to evaluator throughout ([#81](#81)) ([0134682](0134682))
* **evaluators:** split into builtin + extra packages for PyPI ([#5](#5)) ([0e0a78a](0e0a78a))
* **models:** simplify step model and schema ([#70](#70)) ([4c1d637](4c1d637))
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Mar 21, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
nanookclaw added a commit to nanookclaw/agent-control that referenced this pull request Jun 12, 2026
Three issues raised by lan17 in PR review:
1. Float precision on threshold boundary (agentcontrol#1)
baseline=1.0, window=0.9, threshold=0.10: IEEE 754 gives
drift_magnitude=0.09999999... which fails >= 0.10. Fixed with
round(drift_magnitude, 10) >= drift_threshold in _compute_drift().
2. Race condition on concurrent history writes (agentcontrol#3)
load→append→save was not atomic: two workers for the same agent_id
would both read stale history and the last writer would silently drop
the other's observation. Replaced _load_history() / _save_history()
pair with _load_and_append_history() which holds fcntl.LOCK_EX for
the full read-modify-write cycle. Lock is per-agent (.lock file),
so independent agents remain fully parallel.
3. Release wiring missing for drift package (agentcontrol#2)
test-extras, scripts/build.py, Makefile and .PHONY only referenced
galileo. Added drift-{test,lint,lint-fix,typecheck,build} targets to
Makefile, wired drift-test into test-extras, and added
build_evaluator_drift() to scripts/build.py (including 'drift' and
'all' targets).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@namrataghadi-galileo@lan17@nachiket-galileo@abhinav-galileo