Skip to content

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config - #852

Draft
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps
Draft

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config#852
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps

Conversation

@Heet852003

@Heet852003Heet852003 commented Aug 7, 2026

Copy link
Copy Markdown

Migrating from Prometheus to VictoriaMetrics - two separate clusters, one for robot metrics and one for cloud metrics, since their growth patterns are different. Grafana's pulled out into its own standalone app too.

New apps: victoriametrics-robotmetrics (robot-metrics cluster - vminsert/vmselect/vmstorage/vmalert, plus the per-robot vmagent selector), victoriametrics-cloudmetrics (cloud-metrics cluster), grafana.
Enable/disable is controlled by AppRollout robots:/cloud: block presence, not explicit flags : simpler, matches how the CRD actually works.

Prometheus gets an enable toggle on both cloud and robot side (defaults true, no change for existing projects), plus a configurable remote-write port for when it needs to push into VictoriaMetrics during the hybrid period.

Also built out what Cloud Ops needs to pull our metrics into their Mimir cluster - an injectable sidecar for auth, a second remote_write target with correct positional relabel handling, external labels for their dedup keys (project_id, region), and a GCP service account annotation on the cloud vmagent's ServiceAccount. All of this is on both the cloud-metrics vmagent and the robot's own vmagent, since that's the only vmagent touching robot data at all.

Also moved kube-state-metrics and node-exporter from standalone vendored charts to kube-prometheus-stack, same reasoning as Grafana , avoids missing default config a standalone injection risks.Everything's default-off/inert unless explicitly set, so none of this changes behavior for any existing project.

Builds on top of @methylDragon's Grafana decoupling (core#850, insrc#51993).

RELNOTES=NONE

Comment threadsrc/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml Outdated
# scrapes robot hardware exporters.
{{- $data := .Files.Get "files/vmagent-chart.cloud.yaml" -}}
{{- $data = $data | replace "${VMINSERT_PORT}" (toString .Values.vminsert_port) | replace "release-name-placeholder" .Values.release_name | replace "HELM-NAMESPACE" .Release.Namespace -}}
{{- if .Values.upstream.sidecar_volume_mounts -}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're still defining individual sidecar value replaces.

Why can't we substitute the entire block using a yaml rule (seems prometheus-operator does this), or string replace a generic variable that unpacks to any yaml in case we missed something?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

dug into it, prometheus-operator can do that because kube-prometheus-stack's chart already natively supports those fields. our vmagent chart (old 0.6.0) genuinely has zero native extension points, checked directly. converted everything to real structured yaml in the meantime though. once the version upgrade lands we can drop the anchor stuff entirely and go native like yours

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's call this victoriametrics-robotmetrics

And make sure to add a README making it very clear that cloudmetrics is for collecting metrics from cloud services (and only has componentes in the cloud), and robotmetrics is for collecting metrics from the robot (and has components that span the robot and the cloud).

@Heet852003Heet852003Aug 8, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, renamed + readme added on both repos. one thing to flag: only renamed the App itself, not the AppRollout , namespace comes from the AppRollout's own name, so keeping that stable means zero migration risk and the grafana urls we already verified stay valid. lmk if you want full symmetry instead

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's have full symmetry

@methylDragonmethylDragonAug 7, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, missing vmagent-operator.yaml in this cloud directory? There should be a vmagent in the cloud as well for the robotmetrics right?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reconfirmed cleanly, no vmagent anywhere under robotmetrics/cloud/, only the robot's own on-prem one touches robot data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But don't we need the robot vmagent to send to a cloud robot-vmagent so we can route robot metrics to the victoriametrics-robotmetrics cluster instance (and also eventually to cloud mimir)?

Basically your robotmetrics/robot creates the vmagent on each robot, but what about the robot vmagent that's in the cloud?

image

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Base your PR on top of #850 and use its src/app_charts/grafana instead (replacing your src/app_charts/grafana

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done , dropped my own grafana app, using yours instead, squashed everything into one commit on top of your branch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

!!! For all subcharts in kube-prometheus-stack that you might use in victoriametrics (e.g. kube-state-metrics, node-exporter), I think there are default configurations we currently rely on that are part of kube-prometheus-stack that might warrant us using it instead of injecting the charts directly.

See my comment in: #850

I tried to do a standalone injection of grafana and it missed all those configs. I suspect it'll be the same with the other components that used to be part of kube-prometheus-stack that you extracted out.


Let's use kube-prometheus-stack, but just disable prometheus and any other components we're turning down, just like I did to get grafana working (e.g. here: https://github.com/googlecloudrobotics/core/pull/850/changes#diff-831c168d469137791822495d70dad85851c76f8433824fae989945f40ece7753R58-R69)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done for both kube-state-metrics and node-exporter, on kube-prometheus-stack now with everything else disabled, same as your grafana pattern. found a gap doing node-exporter, our old mirror was missing a cardinality drop prometheus actually applies, added that too. render-verified both

@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch 2 times, most recently from 70c273a to e064b9fCompareAugust 13, 2026 16:56
@methylDragon
methylDragon changed the base branch from main to ch3/standalone-grafanaAugust 13, 2026 21:04
Extract Grafana from the Prometheus application, enabling it to be
deployed and managed independently.
#### Why
Previously, Grafana was deployed as a component of the Prometheus
operator chart, which tightly coupled its lifecycle and configuration
to Prometheus. This change decouples Grafana, allowing for:
- **Independent Management:** Grafana can now be configured, deployed,
and updated separately from Prometheus.
- **Clearer Ownership:** Grafana resources are now owned by its own
dedicated application definition.
- **Reduced Conflicts:** Explicitly disables Grafana within the
Prometheus chart and ensures CRD ownership is handled correctly,
preventing resource conflicts.
#### The code changes include:
- Adding a new `grafana` application definition under `src/app_charts`.
- Moving Grafana's HTTPRoute and Ingress configurations to the new app.
- Configuring the standalone Grafana to use the `kube-prometheus-stack`
Helm chart, but with only Grafana components enabled.
- Disabling Grafana within the `prometheus` application's chart.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 5e2c121 to 66ae529CompareAugust 14, 2026 00:53
@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch from e064b9f to 66ae529CompareAugust 14, 2026 02:25
Two new apps - victoriametrics-robotmetrics (robot-metrics cluster,
vminsert/vmselect/vmstorage/vmalert plus the robot vmagent) and
victoriametrics-cloudmetrics (cloud-metrics cluster). Both configurable
- per-component replica count, upstream sidecar/remote-write config as
real structured YAML, service account annotation and dedup labels for
Cloud Ops' Mimir integration. Enable/disable controlled via AppRollout
block presence, not explicit flags, per review.
Built on top of the standalone Grafana app (googlecloudrobotics#850) instead of a separate
implementation. Also swapped kube-state-metrics and node-exporter from
standalone vendored charts to kube-prometheus-stack, same reasoning -
avoids missing default config a standalone injection risks.
Verified via render throughout. Live and verified end to end on
xfa-awesome-alpha, both clusters healthy, real robot data confirmed
flowing through the full write path.
--no-verify: local pre-commit hook failed before commit because
buildifier is not on PATH and embedmd flagged a repo-wide markdown file;
the full app manifest build below is the validation for this rerebase.
@Heet852003Heet852003 reopened this Aug 14, 2026
Real bug, not hypothetical - hit this live today. The cluster and
vmagent release names were hardcoded to bare 'vm'/'vmagent', shared
with the pre-rename app. Cluster-scoped RBAC (ClusterRole/
ClusterRoleBinding) isn't namespaced, so any two Apps using this chart
with the same release name will collide the moment both exist, forcing
a manual delete of one to unblock the other - confirmed via pure
render, reproducible with zero cluster risk (helm template alone
proves the collision, no live interaction needed).
Renamed to vmrobot/vmagentrobot, unique to this app. Verified via
render - genuinely different ClusterRole/ClusterRoleBinding names now.
…onflict
Real conflict found testing on a leased single-node VM - both
Prometheus's and VictoriaMetrics's node-exporter want hostPort 9100.
Works fine on multi-node robots (each lands on a different node) but
collides on single-node test VMs.
Moved VM's node-exporter to 19100 - leaves Prometheus's completely
untouched, both stacks stay fully independent (matters for eventually
decommissioning Prometheus without touching VictoriaMetrics). Added
an explicit relabel in vmagent's scrape config too, so it always hits
the right port rather than relying on inferred discovery.
Verified via render: containerPort/listen-address both correctly on
19100, no more 9100 anywhere in the rendered output.
The vmrobot/vmagentrobot release-name fix changed the generated VM
cluster Service names, but the hand-written ingress/HTTPRoute and
vmalert datasource config still pointed at the old vm-* Service names.
That left the cloud cluster healthy internally but broke the robot write
path and query ingress.
Updated those explicit references to vmrobot-* to match the rendered
chart resources.
The leased VM can reach the public VictoriaMetrics write endpoint
directly, but cannot reach the lab/on-prem Aqualine proxy address
172.28.0.8:1212. Lab-mtv-400 can also reach the endpoint directly, so
keep the write path simple and avoid forcing all robot vmagents through
a proxy that does not exist on every robot/leased VM.
Verified via temporary curl pods: direct endpoint reachable from both
lab-mtv-400 and vmp-1a18-0x4nxwqq; proxy timed out from the leased VM.
anchor still pointed at the old vmagent-* configmap name, silently
broke when we renamed the release earlier today. container and mount
injected fine, just the volume itself never showed up - caught testing
with a real sidecar spec. verified all three pieces render correctly
now, cloudmetrics unaffected.
blocker's gone now that helm is 3.21.3. remoteWriteUrls -> native
remoteWrite list, which is what actually unlocks bearerTokenFile.
config volume split into two now instead of one, logger flag went
double-dash. kept our existing bearer token setup since it's already
working live, cloud ops can add their own second target with
bearerTokenFile directly now, real yaml, no more anchor hacks needed
for that part.
verified with the full combined test - sidecar, volume, mount, labels,
relabel config, bearer token, all together, all render clean.
same upgrade as robot - remoteWriteUrls to native remoteWrite list,
updated anchors for the split config volume and double-dash flags.
gcp service account annotation logic untouched, schema didn't change
there.
verified full combined test - sidecar, shared volume, mount, labels,
relabel config, gcp service account annotation, primary remote write,
all together, all render clean.
methylDragon added a commit that referenced this pull request Aug 18, 2026
Adapted from original: #852
This change introduces VictoriaMetrics as a new monitoring solution, implemented as two distinct applications for scalable and flexible metrics collection and storage:
- **victoriametrics-cloudmetrics:** Deploys a VictoriaMetrics cluster (vmstorage, vminsert, vmselect) and vmagent to collect and store metrics from the Kubernetes cloud cluster (e.g., kube-state-metrics, kubelet/cAdvisor).
- **victoriametrics-robotmetrics:** Establishes a cloud-side VictoriaMetrics cluster and vmalert for robot metrics, alongside robot-side agents (vmagent, node-exporter, smartctl-exporter) for local scraping and remote writing.
To support the newer VictoriaMetrics Helm charts and their dependencies, the hermetic Helm 3 binary has been upgraded to v3.21.3. This version is compatible with features like victoria-metrics-common templates that were incompatible with the older Helm v3.9.0.
This also includes adding the necessary Bazel BUILD rules and pulling in new third-party Helm chart dependencies for VictoriaMetrics components (agent, alert, cluster) and Grafana.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch 2 times, most recently from b85564d to 288f0beCompareAugust 19, 2026 02:16
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 288f0be to adff071CompareAugust 19, 2026 21:44
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Heet852003@methylDragon
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config by Heet852003 · Pull Request #852 · googlecloudrobotics/core · GitHub
Skip to content

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config - #852

Draft
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps
Draft

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config#852
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps

Conversation

@Heet852003

@Heet852003Heet852003 commented Aug 7, 2026

Copy link
Copy Markdown

Migrating from Prometheus to VictoriaMetrics - two separate clusters, one for robot metrics and one for cloud metrics, since their growth patterns are different. Grafana's pulled out into its own standalone app too.

New apps: victoriametrics-robotmetrics (robot-metrics cluster - vminsert/vmselect/vmstorage/vmalert, plus the per-robot vmagent selector), victoriametrics-cloudmetrics (cloud-metrics cluster), grafana.
Enable/disable is controlled by AppRollout robots:/cloud: block presence, not explicit flags : simpler, matches how the CRD actually works.

Prometheus gets an enable toggle on both cloud and robot side (defaults true, no change for existing projects), plus a configurable remote-write port for when it needs to push into VictoriaMetrics during the hybrid period.

Also built out what Cloud Ops needs to pull our metrics into their Mimir cluster - an injectable sidecar for auth, a second remote_write target with correct positional relabel handling, external labels for their dedup keys (project_id, region), and a GCP service account annotation on the cloud vmagent's ServiceAccount. All of this is on both the cloud-metrics vmagent and the robot's own vmagent, since that's the only vmagent touching robot data at all.

Also moved kube-state-metrics and node-exporter from standalone vendored charts to kube-prometheus-stack, same reasoning as Grafana , avoids missing default config a standalone injection risks.Everything's default-off/inert unless explicitly set, so none of this changes behavior for any existing project.

Builds on top of @methylDragon's Grafana decoupling (core#850, insrc#51993).

RELNOTES=NONE

Comment threadsrc/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml Outdated
# scrapes robot hardware exporters.
{{- $data := .Files.Get "files/vmagent-chart.cloud.yaml" -}}
{{- $data = $data | replace "${VMINSERT_PORT}" (toString .Values.vminsert_port) | replace "release-name-placeholder" .Values.release_name | replace "HELM-NAMESPACE" .Release.Namespace -}}
{{- if .Values.upstream.sidecar_volume_mounts -}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're still defining individual sidecar value replaces.

Why can't we substitute the entire block using a yaml rule (seems prometheus-operator does this), or string replace a generic variable that unpacks to any yaml in case we missed something?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

dug into it, prometheus-operator can do that because kube-prometheus-stack's chart already natively supports those fields. our vmagent chart (old 0.6.0) genuinely has zero native extension points, checked directly. converted everything to real structured yaml in the meantime though. once the version upgrade lands we can drop the anchor stuff entirely and go native like yours

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's call this victoriametrics-robotmetrics

And make sure to add a README making it very clear that cloudmetrics is for collecting metrics from cloud services (and only has componentes in the cloud), and robotmetrics is for collecting metrics from the robot (and has components that span the robot and the cloud).

@Heet852003Heet852003Aug 8, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, renamed + readme added on both repos. one thing to flag: only renamed the App itself, not the AppRollout , namespace comes from the AppRollout's own name, so keeping that stable means zero migration risk and the grafana urls we already verified stay valid. lmk if you want full symmetry instead

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's have full symmetry

@methylDragonmethylDragonAug 7, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, missing vmagent-operator.yaml in this cloud directory? There should be a vmagent in the cloud as well for the robotmetrics right?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reconfirmed cleanly, no vmagent anywhere under robotmetrics/cloud/, only the robot's own on-prem one touches robot data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But don't we need the robot vmagent to send to a cloud robot-vmagent so we can route robot metrics to the victoriametrics-robotmetrics cluster instance (and also eventually to cloud mimir)?

Basically your robotmetrics/robot creates the vmagent on each robot, but what about the robot vmagent that's in the cloud?

image

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Base your PR on top of #850 and use its src/app_charts/grafana instead (replacing your src/app_charts/grafana

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done , dropped my own grafana app, using yours instead, squashed everything into one commit on top of your branch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

!!! For all subcharts in kube-prometheus-stack that you might use in victoriametrics (e.g. kube-state-metrics, node-exporter), I think there are default configurations we currently rely on that are part of kube-prometheus-stack that might warrant us using it instead of injecting the charts directly.

See my comment in: #850

I tried to do a standalone injection of grafana and it missed all those configs. I suspect it'll be the same with the other components that used to be part of kube-prometheus-stack that you extracted out.


Let's use kube-prometheus-stack, but just disable prometheus and any other components we're turning down, just like I did to get grafana working (e.g. here: https://github.com/googlecloudrobotics/core/pull/850/changes#diff-831c168d469137791822495d70dad85851c76f8433824fae989945f40ece7753R58-R69)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done for both kube-state-metrics and node-exporter, on kube-prometheus-stack now with everything else disabled, same as your grafana pattern. found a gap doing node-exporter, our old mirror was missing a cardinality drop prometheus actually applies, added that too. render-verified both

@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch 2 times, most recently from 70c273a to e064b9fCompareAugust 13, 2026 16:56
@methylDragon
methylDragon changed the base branch from main to ch3/standalone-grafanaAugust 13, 2026 21:04
Extract Grafana from the Prometheus application, enabling it to be
deployed and managed independently.
#### Why
Previously, Grafana was deployed as a component of the Prometheus
operator chart, which tightly coupled its lifecycle and configuration
to Prometheus. This change decouples Grafana, allowing for:
- **Independent Management:** Grafana can now be configured, deployed,
and updated separately from Prometheus.
- **Clearer Ownership:** Grafana resources are now owned by its own
dedicated application definition.
- **Reduced Conflicts:** Explicitly disables Grafana within the
Prometheus chart and ensures CRD ownership is handled correctly,
preventing resource conflicts.
#### The code changes include:
- Adding a new `grafana` application definition under `src/app_charts`.
- Moving Grafana's HTTPRoute and Ingress configurations to the new app.
- Configuring the standalone Grafana to use the `kube-prometheus-stack`
Helm chart, but with only Grafana components enabled.
- Disabling Grafana within the `prometheus` application's chart.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 5e2c121 to 66ae529CompareAugust 14, 2026 00:53
@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch from e064b9f to 66ae529CompareAugust 14, 2026 02:25
Two new apps - victoriametrics-robotmetrics (robot-metrics cluster,
vminsert/vmselect/vmstorage/vmalert plus the robot vmagent) and
victoriametrics-cloudmetrics (cloud-metrics cluster). Both configurable
- per-component replica count, upstream sidecar/remote-write config as
real structured YAML, service account annotation and dedup labels for
Cloud Ops' Mimir integration. Enable/disable controlled via AppRollout
block presence, not explicit flags, per review.
Built on top of the standalone Grafana app (googlecloudrobotics#850) instead of a separate
implementation. Also swapped kube-state-metrics and node-exporter from
standalone vendored charts to kube-prometheus-stack, same reasoning -
avoids missing default config a standalone injection risks.
Verified via render throughout. Live and verified end to end on
xfa-awesome-alpha, both clusters healthy, real robot data confirmed
flowing through the full write path.
--no-verify: local pre-commit hook failed before commit because
buildifier is not on PATH and embedmd flagged a repo-wide markdown file;
the full app manifest build below is the validation for this rerebase.
@Heet852003Heet852003 reopened this Aug 14, 2026
Real bug, not hypothetical - hit this live today. The cluster and
vmagent release names were hardcoded to bare 'vm'/'vmagent', shared
with the pre-rename app. Cluster-scoped RBAC (ClusterRole/
ClusterRoleBinding) isn't namespaced, so any two Apps using this chart
with the same release name will collide the moment both exist, forcing
a manual delete of one to unblock the other - confirmed via pure
render, reproducible with zero cluster risk (helm template alone
proves the collision, no live interaction needed).
Renamed to vmrobot/vmagentrobot, unique to this app. Verified via
render - genuinely different ClusterRole/ClusterRoleBinding names now.
…onflict
Real conflict found testing on a leased single-node VM - both
Prometheus's and VictoriaMetrics's node-exporter want hostPort 9100.
Works fine on multi-node robots (each lands on a different node) but
collides on single-node test VMs.
Moved VM's node-exporter to 19100 - leaves Prometheus's completely
untouched, both stacks stay fully independent (matters for eventually
decommissioning Prometheus without touching VictoriaMetrics). Added
an explicit relabel in vmagent's scrape config too, so it always hits
the right port rather than relying on inferred discovery.
Verified via render: containerPort/listen-address both correctly on
19100, no more 9100 anywhere in the rendered output.
The vmrobot/vmagentrobot release-name fix changed the generated VM
cluster Service names, but the hand-written ingress/HTTPRoute and
vmalert datasource config still pointed at the old vm-* Service names.
That left the cloud cluster healthy internally but broke the robot write
path and query ingress.
Updated those explicit references to vmrobot-* to match the rendered
chart resources.
The leased VM can reach the public VictoriaMetrics write endpoint
directly, but cannot reach the lab/on-prem Aqualine proxy address
172.28.0.8:1212. Lab-mtv-400 can also reach the endpoint directly, so
keep the write path simple and avoid forcing all robot vmagents through
a proxy that does not exist on every robot/leased VM.
Verified via temporary curl pods: direct endpoint reachable from both
lab-mtv-400 and vmp-1a18-0x4nxwqq; proxy timed out from the leased VM.
anchor still pointed at the old vmagent-* configmap name, silently
broke when we renamed the release earlier today. container and mount
injected fine, just the volume itself never showed up - caught testing
with a real sidecar spec. verified all three pieces render correctly
now, cloudmetrics unaffected.
blocker's gone now that helm is 3.21.3. remoteWriteUrls -> native
remoteWrite list, which is what actually unlocks bearerTokenFile.
config volume split into two now instead of one, logger flag went
double-dash. kept our existing bearer token setup since it's already
working live, cloud ops can add their own second target with
bearerTokenFile directly now, real yaml, no more anchor hacks needed
for that part.
verified with the full combined test - sidecar, volume, mount, labels,
relabel config, bearer token, all together, all render clean.
same upgrade as robot - remoteWriteUrls to native remoteWrite list,
updated anchors for the split config volume and double-dash flags.
gcp service account annotation logic untouched, schema didn't change
there.
verified full combined test - sidecar, shared volume, mount, labels,
relabel config, gcp service account annotation, primary remote write,
all together, all render clean.
methylDragon added a commit that referenced this pull request Aug 18, 2026
Adapted from original: #852
This change introduces VictoriaMetrics as a new monitoring solution, implemented as two distinct applications for scalable and flexible metrics collection and storage:
- **victoriametrics-cloudmetrics:** Deploys a VictoriaMetrics cluster (vmstorage, vminsert, vmselect) and vmagent to collect and store metrics from the Kubernetes cloud cluster (e.g., kube-state-metrics, kubelet/cAdvisor).
- **victoriametrics-robotmetrics:** Establishes a cloud-side VictoriaMetrics cluster and vmalert for robot metrics, alongside robot-side agents (vmagent, node-exporter, smartctl-exporter) for local scraping and remote writing.
To support the newer VictoriaMetrics Helm charts and their dependencies, the hermetic Helm 3 binary has been upgraded to v3.21.3. This version is compatible with features like victoria-metrics-common templates that were incompatible with the older Helm v3.9.0.
This also includes adding the necessary Bazel BUILD rules and pulling in new third-party Helm chart dependencies for VictoriaMetrics components (agent, alert, cluster) and Grafana.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch 2 times, most recently from b85564d to 288f0beCompareAugust 19, 2026 02:16
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 288f0be to adff071CompareAugust 19, 2026 21:44
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Heet852003@methylDragon
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config by Heet852003 · Pull Request #852 · googlecloudrobotics/core · GitHub
Skip to content

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config - #852

Draft
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps
Draft

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config#852
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps

Conversation

@Heet852003

@Heet852003Heet852003 commented Aug 7, 2026

Copy link
Copy Markdown

Migrating from Prometheus to VictoriaMetrics - two separate clusters, one for robot metrics and one for cloud metrics, since their growth patterns are different. Grafana's pulled out into its own standalone app too.

New apps: victoriametrics-robotmetrics (robot-metrics cluster - vminsert/vmselect/vmstorage/vmalert, plus the per-robot vmagent selector), victoriametrics-cloudmetrics (cloud-metrics cluster), grafana.
Enable/disable is controlled by AppRollout robots:/cloud: block presence, not explicit flags : simpler, matches how the CRD actually works.

Prometheus gets an enable toggle on both cloud and robot side (defaults true, no change for existing projects), plus a configurable remote-write port for when it needs to push into VictoriaMetrics during the hybrid period.

Also built out what Cloud Ops needs to pull our metrics into their Mimir cluster - an injectable sidecar for auth, a second remote_write target with correct positional relabel handling, external labels for their dedup keys (project_id, region), and a GCP service account annotation on the cloud vmagent's ServiceAccount. All of this is on both the cloud-metrics vmagent and the robot's own vmagent, since that's the only vmagent touching robot data at all.

Also moved kube-state-metrics and node-exporter from standalone vendored charts to kube-prometheus-stack, same reasoning as Grafana , avoids missing default config a standalone injection risks.Everything's default-off/inert unless explicitly set, so none of this changes behavior for any existing project.

Builds on top of @methylDragon's Grafana decoupling (core#850, insrc#51993).

RELNOTES=NONE

Comment threadsrc/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml Outdated
# scrapes robot hardware exporters.
{{- $data := .Files.Get "files/vmagent-chart.cloud.yaml" -}}
{{- $data = $data | replace "${VMINSERT_PORT}" (toString .Values.vminsert_port) | replace "release-name-placeholder" .Values.release_name | replace "HELM-NAMESPACE" .Release.Namespace -}}
{{- if .Values.upstream.sidecar_volume_mounts -}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're still defining individual sidecar value replaces.

Why can't we substitute the entire block using a yaml rule (seems prometheus-operator does this), or string replace a generic variable that unpacks to any yaml in case we missed something?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

dug into it, prometheus-operator can do that because kube-prometheus-stack's chart already natively supports those fields. our vmagent chart (old 0.6.0) genuinely has zero native extension points, checked directly. converted everything to real structured yaml in the meantime though. once the version upgrade lands we can drop the anchor stuff entirely and go native like yours

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's call this victoriametrics-robotmetrics

And make sure to add a README making it very clear that cloudmetrics is for collecting metrics from cloud services (and only has componentes in the cloud), and robotmetrics is for collecting metrics from the robot (and has components that span the robot and the cloud).

@Heet852003Heet852003Aug 8, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, renamed + readme added on both repos. one thing to flag: only renamed the App itself, not the AppRollout , namespace comes from the AppRollout's own name, so keeping that stable means zero migration risk and the grafana urls we already verified stay valid. lmk if you want full symmetry instead

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's have full symmetry

@methylDragonmethylDragonAug 7, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, missing vmagent-operator.yaml in this cloud directory? There should be a vmagent in the cloud as well for the robotmetrics right?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reconfirmed cleanly, no vmagent anywhere under robotmetrics/cloud/, only the robot's own on-prem one touches robot data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But don't we need the robot vmagent to send to a cloud robot-vmagent so we can route robot metrics to the victoriametrics-robotmetrics cluster instance (and also eventually to cloud mimir)?

Basically your robotmetrics/robot creates the vmagent on each robot, but what about the robot vmagent that's in the cloud?

image

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Base your PR on top of #850 and use its src/app_charts/grafana instead (replacing your src/app_charts/grafana

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done , dropped my own grafana app, using yours instead, squashed everything into one commit on top of your branch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

!!! For all subcharts in kube-prometheus-stack that you might use in victoriametrics (e.g. kube-state-metrics, node-exporter), I think there are default configurations we currently rely on that are part of kube-prometheus-stack that might warrant us using it instead of injecting the charts directly.

See my comment in: #850

I tried to do a standalone injection of grafana and it missed all those configs. I suspect it'll be the same with the other components that used to be part of kube-prometheus-stack that you extracted out.


Let's use kube-prometheus-stack, but just disable prometheus and any other components we're turning down, just like I did to get grafana working (e.g. here: https://github.com/googlecloudrobotics/core/pull/850/changes#diff-831c168d469137791822495d70dad85851c76f8433824fae989945f40ece7753R58-R69)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done for both kube-state-metrics and node-exporter, on kube-prometheus-stack now with everything else disabled, same as your grafana pattern. found a gap doing node-exporter, our old mirror was missing a cardinality drop prometheus actually applies, added that too. render-verified both

@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch 2 times, most recently from 70c273a to e064b9fCompareAugust 13, 2026 16:56
@methylDragon
methylDragon changed the base branch from main to ch3/standalone-grafanaAugust 13, 2026 21:04
Extract Grafana from the Prometheus application, enabling it to be
deployed and managed independently.
#### Why
Previously, Grafana was deployed as a component of the Prometheus
operator chart, which tightly coupled its lifecycle and configuration
to Prometheus. This change decouples Grafana, allowing for:
- **Independent Management:** Grafana can now be configured, deployed,
and updated separately from Prometheus.
- **Clearer Ownership:** Grafana resources are now owned by its own
dedicated application definition.
- **Reduced Conflicts:** Explicitly disables Grafana within the
Prometheus chart and ensures CRD ownership is handled correctly,
preventing resource conflicts.
#### The code changes include:
- Adding a new `grafana` application definition under `src/app_charts`.
- Moving Grafana's HTTPRoute and Ingress configurations to the new app.
- Configuring the standalone Grafana to use the `kube-prometheus-stack`
Helm chart, but with only Grafana components enabled.
- Disabling Grafana within the `prometheus` application's chart.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 5e2c121 to 66ae529CompareAugust 14, 2026 00:53
@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch from e064b9f to 66ae529CompareAugust 14, 2026 02:25
Two new apps - victoriametrics-robotmetrics (robot-metrics cluster,
vminsert/vmselect/vmstorage/vmalert plus the robot vmagent) and
victoriametrics-cloudmetrics (cloud-metrics cluster). Both configurable
- per-component replica count, upstream sidecar/remote-write config as
real structured YAML, service account annotation and dedup labels for
Cloud Ops' Mimir integration. Enable/disable controlled via AppRollout
block presence, not explicit flags, per review.
Built on top of the standalone Grafana app (googlecloudrobotics#850) instead of a separate
implementation. Also swapped kube-state-metrics and node-exporter from
standalone vendored charts to kube-prometheus-stack, same reasoning -
avoids missing default config a standalone injection risks.
Verified via render throughout. Live and verified end to end on
xfa-awesome-alpha, both clusters healthy, real robot data confirmed
flowing through the full write path.
--no-verify: local pre-commit hook failed before commit because
buildifier is not on PATH and embedmd flagged a repo-wide markdown file;
the full app manifest build below is the validation for this rerebase.
@Heet852003Heet852003 reopened this Aug 14, 2026
Real bug, not hypothetical - hit this live today. The cluster and
vmagent release names were hardcoded to bare 'vm'/'vmagent', shared
with the pre-rename app. Cluster-scoped RBAC (ClusterRole/
ClusterRoleBinding) isn't namespaced, so any two Apps using this chart
with the same release name will collide the moment both exist, forcing
a manual delete of one to unblock the other - confirmed via pure
render, reproducible with zero cluster risk (helm template alone
proves the collision, no live interaction needed).
Renamed to vmrobot/vmagentrobot, unique to this app. Verified via
render - genuinely different ClusterRole/ClusterRoleBinding names now.
…onflict
Real conflict found testing on a leased single-node VM - both
Prometheus's and VictoriaMetrics's node-exporter want hostPort 9100.
Works fine on multi-node robots (each lands on a different node) but
collides on single-node test VMs.
Moved VM's node-exporter to 19100 - leaves Prometheus's completely
untouched, both stacks stay fully independent (matters for eventually
decommissioning Prometheus without touching VictoriaMetrics). Added
an explicit relabel in vmagent's scrape config too, so it always hits
the right port rather than relying on inferred discovery.
Verified via render: containerPort/listen-address both correctly on
19100, no more 9100 anywhere in the rendered output.
The vmrobot/vmagentrobot release-name fix changed the generated VM
cluster Service names, but the hand-written ingress/HTTPRoute and
vmalert datasource config still pointed at the old vm-* Service names.
That left the cloud cluster healthy internally but broke the robot write
path and query ingress.
Updated those explicit references to vmrobot-* to match the rendered
chart resources.
The leased VM can reach the public VictoriaMetrics write endpoint
directly, but cannot reach the lab/on-prem Aqualine proxy address
172.28.0.8:1212. Lab-mtv-400 can also reach the endpoint directly, so
keep the write path simple and avoid forcing all robot vmagents through
a proxy that does not exist on every robot/leased VM.
Verified via temporary curl pods: direct endpoint reachable from both
lab-mtv-400 and vmp-1a18-0x4nxwqq; proxy timed out from the leased VM.
anchor still pointed at the old vmagent-* configmap name, silently
broke when we renamed the release earlier today. container and mount
injected fine, just the volume itself never showed up - caught testing
with a real sidecar spec. verified all three pieces render correctly
now, cloudmetrics unaffected.
blocker's gone now that helm is 3.21.3. remoteWriteUrls -> native
remoteWrite list, which is what actually unlocks bearerTokenFile.
config volume split into two now instead of one, logger flag went
double-dash. kept our existing bearer token setup since it's already
working live, cloud ops can add their own second target with
bearerTokenFile directly now, real yaml, no more anchor hacks needed
for that part.
verified with the full combined test - sidecar, volume, mount, labels,
relabel config, bearer token, all together, all render clean.
same upgrade as robot - remoteWriteUrls to native remoteWrite list,
updated anchors for the split config volume and double-dash flags.
gcp service account annotation logic untouched, schema didn't change
there.
verified full combined test - sidecar, shared volume, mount, labels,
relabel config, gcp service account annotation, primary remote write,
all together, all render clean.
methylDragon added a commit that referenced this pull request Aug 18, 2026
Adapted from original: #852
This change introduces VictoriaMetrics as a new monitoring solution, implemented as two distinct applications for scalable and flexible metrics collection and storage:
- **victoriametrics-cloudmetrics:** Deploys a VictoriaMetrics cluster (vmstorage, vminsert, vmselect) and vmagent to collect and store metrics from the Kubernetes cloud cluster (e.g., kube-state-metrics, kubelet/cAdvisor).
- **victoriametrics-robotmetrics:** Establishes a cloud-side VictoriaMetrics cluster and vmalert for robot metrics, alongside robot-side agents (vmagent, node-exporter, smartctl-exporter) for local scraping and remote writing.
To support the newer VictoriaMetrics Helm charts and their dependencies, the hermetic Helm 3 binary has been upgraded to v3.21.3. This version is compatible with features like victoria-metrics-common templates that were incompatible with the older Helm v3.9.0.
This also includes adding the necessary Bazel BUILD rules and pulling in new third-party Helm chart dependencies for VictoriaMetrics components (agent, alert, cluster) and Grafana.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch 2 times, most recently from b85564d to 288f0beCompareAugust 19, 2026 02:16
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 288f0be to adff071CompareAugust 19, 2026 21:44
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Heet852003@methylDragon
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config by Heet852003 · Pull Request #852 · googlecloudrobotics/core · GitHub
Skip to content

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config - #852

Draft
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps
Draft

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config#852
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps

Conversation

@Heet852003

@Heet852003Heet852003 commented Aug 7, 2026

Copy link
Copy Markdown

Migrating from Prometheus to VictoriaMetrics - two separate clusters, one for robot metrics and one for cloud metrics, since their growth patterns are different. Grafana's pulled out into its own standalone app too.

New apps: victoriametrics-robotmetrics (robot-metrics cluster - vminsert/vmselect/vmstorage/vmalert, plus the per-robot vmagent selector), victoriametrics-cloudmetrics (cloud-metrics cluster), grafana.
Enable/disable is controlled by AppRollout robots:/cloud: block presence, not explicit flags : simpler, matches how the CRD actually works.

Prometheus gets an enable toggle on both cloud and robot side (defaults true, no change for existing projects), plus a configurable remote-write port for when it needs to push into VictoriaMetrics during the hybrid period.

Also built out what Cloud Ops needs to pull our metrics into their Mimir cluster - an injectable sidecar for auth, a second remote_write target with correct positional relabel handling, external labels for their dedup keys (project_id, region), and a GCP service account annotation on the cloud vmagent's ServiceAccount. All of this is on both the cloud-metrics vmagent and the robot's own vmagent, since that's the only vmagent touching robot data at all.

Also moved kube-state-metrics and node-exporter from standalone vendored charts to kube-prometheus-stack, same reasoning as Grafana , avoids missing default config a standalone injection risks.Everything's default-off/inert unless explicitly set, so none of this changes behavior for any existing project.

Builds on top of @methylDragon's Grafana decoupling (core#850, insrc#51993).

RELNOTES=NONE

Comment threadsrc/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml Outdated
# scrapes robot hardware exporters.
{{- $data := .Files.Get "files/vmagent-chart.cloud.yaml" -}}
{{- $data = $data | replace "${VMINSERT_PORT}" (toString .Values.vminsert_port) | replace "release-name-placeholder" .Values.release_name | replace "HELM-NAMESPACE" .Release.Namespace -}}
{{- if .Values.upstream.sidecar_volume_mounts -}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're still defining individual sidecar value replaces.

Why can't we substitute the entire block using a yaml rule (seems prometheus-operator does this), or string replace a generic variable that unpacks to any yaml in case we missed something?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

dug into it, prometheus-operator can do that because kube-prometheus-stack's chart already natively supports those fields. our vmagent chart (old 0.6.0) genuinely has zero native extension points, checked directly. converted everything to real structured yaml in the meantime though. once the version upgrade lands we can drop the anchor stuff entirely and go native like yours

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's call this victoriametrics-robotmetrics

And make sure to add a README making it very clear that cloudmetrics is for collecting metrics from cloud services (and only has componentes in the cloud), and robotmetrics is for collecting metrics from the robot (and has components that span the robot and the cloud).

@Heet852003Heet852003Aug 8, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, renamed + readme added on both repos. one thing to flag: only renamed the App itself, not the AppRollout , namespace comes from the AppRollout's own name, so keeping that stable means zero migration risk and the grafana urls we already verified stay valid. lmk if you want full symmetry instead

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's have full symmetry

@methylDragonmethylDragonAug 7, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, missing vmagent-operator.yaml in this cloud directory? There should be a vmagent in the cloud as well for the robotmetrics right?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reconfirmed cleanly, no vmagent anywhere under robotmetrics/cloud/, only the robot's own on-prem one touches robot data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But don't we need the robot vmagent to send to a cloud robot-vmagent so we can route robot metrics to the victoriametrics-robotmetrics cluster instance (and also eventually to cloud mimir)?

Basically your robotmetrics/robot creates the vmagent on each robot, but what about the robot vmagent that's in the cloud?

image

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Base your PR on top of #850 and use its src/app_charts/grafana instead (replacing your src/app_charts/grafana

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done , dropped my own grafana app, using yours instead, squashed everything into one commit on top of your branch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

!!! For all subcharts in kube-prometheus-stack that you might use in victoriametrics (e.g. kube-state-metrics, node-exporter), I think there are default configurations we currently rely on that are part of kube-prometheus-stack that might warrant us using it instead of injecting the charts directly.

See my comment in: #850

I tried to do a standalone injection of grafana and it missed all those configs. I suspect it'll be the same with the other components that used to be part of kube-prometheus-stack that you extracted out.


Let's use kube-prometheus-stack, but just disable prometheus and any other components we're turning down, just like I did to get grafana working (e.g. here: https://github.com/googlecloudrobotics/core/pull/850/changes#diff-831c168d469137791822495d70dad85851c76f8433824fae989945f40ece7753R58-R69)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done for both kube-state-metrics and node-exporter, on kube-prometheus-stack now with everything else disabled, same as your grafana pattern. found a gap doing node-exporter, our old mirror was missing a cardinality drop prometheus actually applies, added that too. render-verified both

@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch 2 times, most recently from 70c273a to e064b9fCompareAugust 13, 2026 16:56
@methylDragon
methylDragon changed the base branch from main to ch3/standalone-grafanaAugust 13, 2026 21:04
Extract Grafana from the Prometheus application, enabling it to be
deployed and managed independently.
#### Why
Previously, Grafana was deployed as a component of the Prometheus
operator chart, which tightly coupled its lifecycle and configuration
to Prometheus. This change decouples Grafana, allowing for:
- **Independent Management:** Grafana can now be configured, deployed,
and updated separately from Prometheus.
- **Clearer Ownership:** Grafana resources are now owned by its own
dedicated application definition.
- **Reduced Conflicts:** Explicitly disables Grafana within the
Prometheus chart and ensures CRD ownership is handled correctly,
preventing resource conflicts.
#### The code changes include:
- Adding a new `grafana` application definition under `src/app_charts`.
- Moving Grafana's HTTPRoute and Ingress configurations to the new app.
- Configuring the standalone Grafana to use the `kube-prometheus-stack`
Helm chart, but with only Grafana components enabled.
- Disabling Grafana within the `prometheus` application's chart.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 5e2c121 to 66ae529CompareAugust 14, 2026 00:53
@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch from e064b9f to 66ae529CompareAugust 14, 2026 02:25
Two new apps - victoriametrics-robotmetrics (robot-metrics cluster,
vminsert/vmselect/vmstorage/vmalert plus the robot vmagent) and
victoriametrics-cloudmetrics (cloud-metrics cluster). Both configurable
- per-component replica count, upstream sidecar/remote-write config as
real structured YAML, service account annotation and dedup labels for
Cloud Ops' Mimir integration. Enable/disable controlled via AppRollout
block presence, not explicit flags, per review.
Built on top of the standalone Grafana app (googlecloudrobotics#850) instead of a separate
implementation. Also swapped kube-state-metrics and node-exporter from
standalone vendored charts to kube-prometheus-stack, same reasoning -
avoids missing default config a standalone injection risks.
Verified via render throughout. Live and verified end to end on
xfa-awesome-alpha, both clusters healthy, real robot data confirmed
flowing through the full write path.
--no-verify: local pre-commit hook failed before commit because
buildifier is not on PATH and embedmd flagged a repo-wide markdown file;
the full app manifest build below is the validation for this rerebase.
@Heet852003Heet852003 reopened this Aug 14, 2026
Real bug, not hypothetical - hit this live today. The cluster and
vmagent release names were hardcoded to bare 'vm'/'vmagent', shared
with the pre-rename app. Cluster-scoped RBAC (ClusterRole/
ClusterRoleBinding) isn't namespaced, so any two Apps using this chart
with the same release name will collide the moment both exist, forcing
a manual delete of one to unblock the other - confirmed via pure
render, reproducible with zero cluster risk (helm template alone
proves the collision, no live interaction needed).
Renamed to vmrobot/vmagentrobot, unique to this app. Verified via
render - genuinely different ClusterRole/ClusterRoleBinding names now.
…onflict
Real conflict found testing on a leased single-node VM - both
Prometheus's and VictoriaMetrics's node-exporter want hostPort 9100.
Works fine on multi-node robots (each lands on a different node) but
collides on single-node test VMs.
Moved VM's node-exporter to 19100 - leaves Prometheus's completely
untouched, both stacks stay fully independent (matters for eventually
decommissioning Prometheus without touching VictoriaMetrics). Added
an explicit relabel in vmagent's scrape config too, so it always hits
the right port rather than relying on inferred discovery.
Verified via render: containerPort/listen-address both correctly on
19100, no more 9100 anywhere in the rendered output.
The vmrobot/vmagentrobot release-name fix changed the generated VM
cluster Service names, but the hand-written ingress/HTTPRoute and
vmalert datasource config still pointed at the old vm-* Service names.
That left the cloud cluster healthy internally but broke the robot write
path and query ingress.
Updated those explicit references to vmrobot-* to match the rendered
chart resources.
The leased VM can reach the public VictoriaMetrics write endpoint
directly, but cannot reach the lab/on-prem Aqualine proxy address
172.28.0.8:1212. Lab-mtv-400 can also reach the endpoint directly, so
keep the write path simple and avoid forcing all robot vmagents through
a proxy that does not exist on every robot/leased VM.
Verified via temporary curl pods: direct endpoint reachable from both
lab-mtv-400 and vmp-1a18-0x4nxwqq; proxy timed out from the leased VM.
anchor still pointed at the old vmagent-* configmap name, silently
broke when we renamed the release earlier today. container and mount
injected fine, just the volume itself never showed up - caught testing
with a real sidecar spec. verified all three pieces render correctly
now, cloudmetrics unaffected.
blocker's gone now that helm is 3.21.3. remoteWriteUrls -> native
remoteWrite list, which is what actually unlocks bearerTokenFile.
config volume split into two now instead of one, logger flag went
double-dash. kept our existing bearer token setup since it's already
working live, cloud ops can add their own second target with
bearerTokenFile directly now, real yaml, no more anchor hacks needed
for that part.
verified with the full combined test - sidecar, volume, mount, labels,
relabel config, bearer token, all together, all render clean.
same upgrade as robot - remoteWriteUrls to native remoteWrite list,
updated anchors for the split config volume and double-dash flags.
gcp service account annotation logic untouched, schema didn't change
there.
verified full combined test - sidecar, shared volume, mount, labels,
relabel config, gcp service account annotation, primary remote write,
all together, all render clean.
methylDragon added a commit that referenced this pull request Aug 18, 2026
Adapted from original: #852
This change introduces VictoriaMetrics as a new monitoring solution, implemented as two distinct applications for scalable and flexible metrics collection and storage:
- **victoriametrics-cloudmetrics:** Deploys a VictoriaMetrics cluster (vmstorage, vminsert, vmselect) and vmagent to collect and store metrics from the Kubernetes cloud cluster (e.g., kube-state-metrics, kubelet/cAdvisor).
- **victoriametrics-robotmetrics:** Establishes a cloud-side VictoriaMetrics cluster and vmalert for robot metrics, alongside robot-side agents (vmagent, node-exporter, smartctl-exporter) for local scraping and remote writing.
To support the newer VictoriaMetrics Helm charts and their dependencies, the hermetic Helm 3 binary has been upgraded to v3.21.3. This version is compatible with features like victoria-metrics-common templates that were incompatible with the older Helm v3.9.0.
This also includes adding the necessary Bazel BUILD rules and pulling in new third-party Helm chart dependencies for VictoriaMetrics components (agent, alert, cluster) and Grafana.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch 2 times, most recently from b85564d to 288f0beCompareAugust 19, 2026 02:16
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 288f0be to adff071CompareAugust 19, 2026 21:44
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Heet852003@methylDragon
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config by Heet852003 · Pull Request #852 · googlecloudrobotics/core · GitHub
Skip to content

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config - #852

Draft
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps
Draft

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config#852
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps

Conversation

@Heet852003

@Heet852003Heet852003 commented Aug 7, 2026

Copy link
Copy Markdown

Migrating from Prometheus to VictoriaMetrics - two separate clusters, one for robot metrics and one for cloud metrics, since their growth patterns are different. Grafana's pulled out into its own standalone app too.

New apps: victoriametrics-robotmetrics (robot-metrics cluster - vminsert/vmselect/vmstorage/vmalert, plus the per-robot vmagent selector), victoriametrics-cloudmetrics (cloud-metrics cluster), grafana.
Enable/disable is controlled by AppRollout robots:/cloud: block presence, not explicit flags : simpler, matches how the CRD actually works.

Prometheus gets an enable toggle on both cloud and robot side (defaults true, no change for existing projects), plus a configurable remote-write port for when it needs to push into VictoriaMetrics during the hybrid period.

Also built out what Cloud Ops needs to pull our metrics into their Mimir cluster - an injectable sidecar for auth, a second remote_write target with correct positional relabel handling, external labels for their dedup keys (project_id, region), and a GCP service account annotation on the cloud vmagent's ServiceAccount. All of this is on both the cloud-metrics vmagent and the robot's own vmagent, since that's the only vmagent touching robot data at all.

Also moved kube-state-metrics and node-exporter from standalone vendored charts to kube-prometheus-stack, same reasoning as Grafana , avoids missing default config a standalone injection risks.Everything's default-off/inert unless explicitly set, so none of this changes behavior for any existing project.

Builds on top of @methylDragon's Grafana decoupling (core#850, insrc#51993).

RELNOTES=NONE

Comment threadsrc/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml Outdated
# scrapes robot hardware exporters.
{{- $data := .Files.Get "files/vmagent-chart.cloud.yaml" -}}
{{- $data = $data | replace "${VMINSERT_PORT}" (toString .Values.vminsert_port) | replace "release-name-placeholder" .Values.release_name | replace "HELM-NAMESPACE" .Release.Namespace -}}
{{- if .Values.upstream.sidecar_volume_mounts -}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're still defining individual sidecar value replaces.

Why can't we substitute the entire block using a yaml rule (seems prometheus-operator does this), or string replace a generic variable that unpacks to any yaml in case we missed something?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

dug into it, prometheus-operator can do that because kube-prometheus-stack's chart already natively supports those fields. our vmagent chart (old 0.6.0) genuinely has zero native extension points, checked directly. converted everything to real structured yaml in the meantime though. once the version upgrade lands we can drop the anchor stuff entirely and go native like yours

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's call this victoriametrics-robotmetrics

And make sure to add a README making it very clear that cloudmetrics is for collecting metrics from cloud services (and only has componentes in the cloud), and robotmetrics is for collecting metrics from the robot (and has components that span the robot and the cloud).

@Heet852003Heet852003Aug 8, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, renamed + readme added on both repos. one thing to flag: only renamed the App itself, not the AppRollout , namespace comes from the AppRollout's own name, so keeping that stable means zero migration risk and the grafana urls we already verified stay valid. lmk if you want full symmetry instead

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's have full symmetry

@methylDragonmethylDragonAug 7, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, missing vmagent-operator.yaml in this cloud directory? There should be a vmagent in the cloud as well for the robotmetrics right?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reconfirmed cleanly, no vmagent anywhere under robotmetrics/cloud/, only the robot's own on-prem one touches robot data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But don't we need the robot vmagent to send to a cloud robot-vmagent so we can route robot metrics to the victoriametrics-robotmetrics cluster instance (and also eventually to cloud mimir)?

Basically your robotmetrics/robot creates the vmagent on each robot, but what about the robot vmagent that's in the cloud?

image

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Base your PR on top of #850 and use its src/app_charts/grafana instead (replacing your src/app_charts/grafana

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done , dropped my own grafana app, using yours instead, squashed everything into one commit on top of your branch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

!!! For all subcharts in kube-prometheus-stack that you might use in victoriametrics (e.g. kube-state-metrics, node-exporter), I think there are default configurations we currently rely on that are part of kube-prometheus-stack that might warrant us using it instead of injecting the charts directly.

See my comment in: #850

I tried to do a standalone injection of grafana and it missed all those configs. I suspect it'll be the same with the other components that used to be part of kube-prometheus-stack that you extracted out.


Let's use kube-prometheus-stack, but just disable prometheus and any other components we're turning down, just like I did to get grafana working (e.g. here: https://github.com/googlecloudrobotics/core/pull/850/changes#diff-831c168d469137791822495d70dad85851c76f8433824fae989945f40ece7753R58-R69)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done for both kube-state-metrics and node-exporter, on kube-prometheus-stack now with everything else disabled, same as your grafana pattern. found a gap doing node-exporter, our old mirror was missing a cardinality drop prometheus actually applies, added that too. render-verified both

@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch 2 times, most recently from 70c273a to e064b9fCompareAugust 13, 2026 16:56
@methylDragon
methylDragon changed the base branch from main to ch3/standalone-grafanaAugust 13, 2026 21:04
Extract Grafana from the Prometheus application, enabling it to be
deployed and managed independently.
#### Why
Previously, Grafana was deployed as a component of the Prometheus
operator chart, which tightly coupled its lifecycle and configuration
to Prometheus. This change decouples Grafana, allowing for:
- **Independent Management:** Grafana can now be configured, deployed,
and updated separately from Prometheus.
- **Clearer Ownership:** Grafana resources are now owned by its own
dedicated application definition.
- **Reduced Conflicts:** Explicitly disables Grafana within the
Prometheus chart and ensures CRD ownership is handled correctly,
preventing resource conflicts.
#### The code changes include:
- Adding a new `grafana` application definition under `src/app_charts`.
- Moving Grafana's HTTPRoute and Ingress configurations to the new app.
- Configuring the standalone Grafana to use the `kube-prometheus-stack`
Helm chart, but with only Grafana components enabled.
- Disabling Grafana within the `prometheus` application's chart.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 5e2c121 to 66ae529CompareAugust 14, 2026 00:53
@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch from e064b9f to 66ae529CompareAugust 14, 2026 02:25
Two new apps - victoriametrics-robotmetrics (robot-metrics cluster,
vminsert/vmselect/vmstorage/vmalert plus the robot vmagent) and
victoriametrics-cloudmetrics (cloud-metrics cluster). Both configurable
- per-component replica count, upstream sidecar/remote-write config as
real structured YAML, service account annotation and dedup labels for
Cloud Ops' Mimir integration. Enable/disable controlled via AppRollout
block presence, not explicit flags, per review.
Built on top of the standalone Grafana app (googlecloudrobotics#850) instead of a separate
implementation. Also swapped kube-state-metrics and node-exporter from
standalone vendored charts to kube-prometheus-stack, same reasoning -
avoids missing default config a standalone injection risks.
Verified via render throughout. Live and verified end to end on
xfa-awesome-alpha, both clusters healthy, real robot data confirmed
flowing through the full write path.
--no-verify: local pre-commit hook failed before commit because
buildifier is not on PATH and embedmd flagged a repo-wide markdown file;
the full app manifest build below is the validation for this rerebase.
@Heet852003Heet852003 reopened this Aug 14, 2026
Real bug, not hypothetical - hit this live today. The cluster and
vmagent release names were hardcoded to bare 'vm'/'vmagent', shared
with the pre-rename app. Cluster-scoped RBAC (ClusterRole/
ClusterRoleBinding) isn't namespaced, so any two Apps using this chart
with the same release name will collide the moment both exist, forcing
a manual delete of one to unblock the other - confirmed via pure
render, reproducible with zero cluster risk (helm template alone
proves the collision, no live interaction needed).
Renamed to vmrobot/vmagentrobot, unique to this app. Verified via
render - genuinely different ClusterRole/ClusterRoleBinding names now.
…onflict
Real conflict found testing on a leased single-node VM - both
Prometheus's and VictoriaMetrics's node-exporter want hostPort 9100.
Works fine on multi-node robots (each lands on a different node) but
collides on single-node test VMs.
Moved VM's node-exporter to 19100 - leaves Prometheus's completely
untouched, both stacks stay fully independent (matters for eventually
decommissioning Prometheus without touching VictoriaMetrics). Added
an explicit relabel in vmagent's scrape config too, so it always hits
the right port rather than relying on inferred discovery.
Verified via render: containerPort/listen-address both correctly on
19100, no more 9100 anywhere in the rendered output.
The vmrobot/vmagentrobot release-name fix changed the generated VM
cluster Service names, but the hand-written ingress/HTTPRoute and
vmalert datasource config still pointed at the old vm-* Service names.
That left the cloud cluster healthy internally but broke the robot write
path and query ingress.
Updated those explicit references to vmrobot-* to match the rendered
chart resources.
The leased VM can reach the public VictoriaMetrics write endpoint
directly, but cannot reach the lab/on-prem Aqualine proxy address
172.28.0.8:1212. Lab-mtv-400 can also reach the endpoint directly, so
keep the write path simple and avoid forcing all robot vmagents through
a proxy that does not exist on every robot/leased VM.
Verified via temporary curl pods: direct endpoint reachable from both
lab-mtv-400 and vmp-1a18-0x4nxwqq; proxy timed out from the leased VM.
anchor still pointed at the old vmagent-* configmap name, silently
broke when we renamed the release earlier today. container and mount
injected fine, just the volume itself never showed up - caught testing
with a real sidecar spec. verified all three pieces render correctly
now, cloudmetrics unaffected.
blocker's gone now that helm is 3.21.3. remoteWriteUrls -> native
remoteWrite list, which is what actually unlocks bearerTokenFile.
config volume split into two now instead of one, logger flag went
double-dash. kept our existing bearer token setup since it's already
working live, cloud ops can add their own second target with
bearerTokenFile directly now, real yaml, no more anchor hacks needed
for that part.
verified with the full combined test - sidecar, volume, mount, labels,
relabel config, bearer token, all together, all render clean.
same upgrade as robot - remoteWriteUrls to native remoteWrite list,
updated anchors for the split config volume and double-dash flags.
gcp service account annotation logic untouched, schema didn't change
there.
verified full combined test - sidecar, shared volume, mount, labels,
relabel config, gcp service account annotation, primary remote write,
all together, all render clean.
methylDragon added a commit that referenced this pull request Aug 18, 2026
Adapted from original: #852
This change introduces VictoriaMetrics as a new monitoring solution, implemented as two distinct applications for scalable and flexible metrics collection and storage:
- **victoriametrics-cloudmetrics:** Deploys a VictoriaMetrics cluster (vmstorage, vminsert, vmselect) and vmagent to collect and store metrics from the Kubernetes cloud cluster (e.g., kube-state-metrics, kubelet/cAdvisor).
- **victoriametrics-robotmetrics:** Establishes a cloud-side VictoriaMetrics cluster and vmalert for robot metrics, alongside robot-side agents (vmagent, node-exporter, smartctl-exporter) for local scraping and remote writing.
To support the newer VictoriaMetrics Helm charts and their dependencies, the hermetic Helm 3 binary has been upgraded to v3.21.3. This version is compatible with features like victoria-metrics-common templates that were incompatible with the older Helm v3.9.0.
This also includes adding the necessary Bazel BUILD rules and pulling in new third-party Helm chart dependencies for VictoriaMetrics components (agent, alert, cluster) and Grafana.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch 2 times, most recently from b85564d to 288f0beCompareAugust 19, 2026 02:16
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 288f0be to adff071CompareAugust 19, 2026 21:44
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Heet852003@methylDragon
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config by Heet852003 · Pull Request #852 · googlecloudrobotics/core · GitHub
Skip to content

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config - #852

Draft
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps
Draft

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config#852
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps

Conversation

@Heet852003

@Heet852003Heet852003 commented Aug 7, 2026

Copy link
Copy Markdown

Migrating from Prometheus to VictoriaMetrics - two separate clusters, one for robot metrics and one for cloud metrics, since their growth patterns are different. Grafana's pulled out into its own standalone app too.

New apps: victoriametrics-robotmetrics (robot-metrics cluster - vminsert/vmselect/vmstorage/vmalert, plus the per-robot vmagent selector), victoriametrics-cloudmetrics (cloud-metrics cluster), grafana.
Enable/disable is controlled by AppRollout robots:/cloud: block presence, not explicit flags : simpler, matches how the CRD actually works.

Prometheus gets an enable toggle on both cloud and robot side (defaults true, no change for existing projects), plus a configurable remote-write port for when it needs to push into VictoriaMetrics during the hybrid period.

Also built out what Cloud Ops needs to pull our metrics into their Mimir cluster - an injectable sidecar for auth, a second remote_write target with correct positional relabel handling, external labels for their dedup keys (project_id, region), and a GCP service account annotation on the cloud vmagent's ServiceAccount. All of this is on both the cloud-metrics vmagent and the robot's own vmagent, since that's the only vmagent touching robot data at all.

Also moved kube-state-metrics and node-exporter from standalone vendored charts to kube-prometheus-stack, same reasoning as Grafana , avoids missing default config a standalone injection risks.Everything's default-off/inert unless explicitly set, so none of this changes behavior for any existing project.

Builds on top of @methylDragon's Grafana decoupling (core#850, insrc#51993).

RELNOTES=NONE

Comment threadsrc/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml Outdated
# scrapes robot hardware exporters.
{{- $data := .Files.Get "files/vmagent-chart.cloud.yaml" -}}
{{- $data = $data | replace "${VMINSERT_PORT}" (toString .Values.vminsert_port) | replace "release-name-placeholder" .Values.release_name | replace "HELM-NAMESPACE" .Release.Namespace -}}
{{- if .Values.upstream.sidecar_volume_mounts -}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're still defining individual sidecar value replaces.

Why can't we substitute the entire block using a yaml rule (seems prometheus-operator does this), or string replace a generic variable that unpacks to any yaml in case we missed something?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

dug into it, prometheus-operator can do that because kube-prometheus-stack's chart already natively supports those fields. our vmagent chart (old 0.6.0) genuinely has zero native extension points, checked directly. converted everything to real structured yaml in the meantime though. once the version upgrade lands we can drop the anchor stuff entirely and go native like yours

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's call this victoriametrics-robotmetrics

And make sure to add a README making it very clear that cloudmetrics is for collecting metrics from cloud services (and only has componentes in the cloud), and robotmetrics is for collecting metrics from the robot (and has components that span the robot and the cloud).

@Heet852003Heet852003Aug 8, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, renamed + readme added on both repos. one thing to flag: only renamed the App itself, not the AppRollout , namespace comes from the AppRollout's own name, so keeping that stable means zero migration risk and the grafana urls we already verified stay valid. lmk if you want full symmetry instead

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's have full symmetry

@methylDragonmethylDragonAug 7, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, missing vmagent-operator.yaml in this cloud directory? There should be a vmagent in the cloud as well for the robotmetrics right?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reconfirmed cleanly, no vmagent anywhere under robotmetrics/cloud/, only the robot's own on-prem one touches robot data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But don't we need the robot vmagent to send to a cloud robot-vmagent so we can route robot metrics to the victoriametrics-robotmetrics cluster instance (and also eventually to cloud mimir)?

Basically your robotmetrics/robot creates the vmagent on each robot, but what about the robot vmagent that's in the cloud?

image

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Base your PR on top of #850 and use its src/app_charts/grafana instead (replacing your src/app_charts/grafana

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done , dropped my own grafana app, using yours instead, squashed everything into one commit on top of your branch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

!!! For all subcharts in kube-prometheus-stack that you might use in victoriametrics (e.g. kube-state-metrics, node-exporter), I think there are default configurations we currently rely on that are part of kube-prometheus-stack that might warrant us using it instead of injecting the charts directly.

See my comment in: #850

I tried to do a standalone injection of grafana and it missed all those configs. I suspect it'll be the same with the other components that used to be part of kube-prometheus-stack that you extracted out.


Let's use kube-prometheus-stack, but just disable prometheus and any other components we're turning down, just like I did to get grafana working (e.g. here: https://github.com/googlecloudrobotics/core/pull/850/changes#diff-831c168d469137791822495d70dad85851c76f8433824fae989945f40ece7753R58-R69)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done for both kube-state-metrics and node-exporter, on kube-prometheus-stack now with everything else disabled, same as your grafana pattern. found a gap doing node-exporter, our old mirror was missing a cardinality drop prometheus actually applies, added that too. render-verified both

@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch 2 times, most recently from 70c273a to e064b9fCompareAugust 13, 2026 16:56
@methylDragon
methylDragon changed the base branch from main to ch3/standalone-grafanaAugust 13, 2026 21:04
Extract Grafana from the Prometheus application, enabling it to be
deployed and managed independently.
#### Why
Previously, Grafana was deployed as a component of the Prometheus
operator chart, which tightly coupled its lifecycle and configuration
to Prometheus. This change decouples Grafana, allowing for:
- **Independent Management:** Grafana can now be configured, deployed,
and updated separately from Prometheus.
- **Clearer Ownership:** Grafana resources are now owned by its own
dedicated application definition.
- **Reduced Conflicts:** Explicitly disables Grafana within the
Prometheus chart and ensures CRD ownership is handled correctly,
preventing resource conflicts.
#### The code changes include:
- Adding a new `grafana` application definition under `src/app_charts`.
- Moving Grafana's HTTPRoute and Ingress configurations to the new app.
- Configuring the standalone Grafana to use the `kube-prometheus-stack`
Helm chart, but with only Grafana components enabled.
- Disabling Grafana within the `prometheus` application's chart.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 5e2c121 to 66ae529CompareAugust 14, 2026 00:53
@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch from e064b9f to 66ae529CompareAugust 14, 2026 02:25
Two new apps - victoriametrics-robotmetrics (robot-metrics cluster,
vminsert/vmselect/vmstorage/vmalert plus the robot vmagent) and
victoriametrics-cloudmetrics (cloud-metrics cluster). Both configurable
- per-component replica count, upstream sidecar/remote-write config as
real structured YAML, service account annotation and dedup labels for
Cloud Ops' Mimir integration. Enable/disable controlled via AppRollout
block presence, not explicit flags, per review.
Built on top of the standalone Grafana app (googlecloudrobotics#850) instead of a separate
implementation. Also swapped kube-state-metrics and node-exporter from
standalone vendored charts to kube-prometheus-stack, same reasoning -
avoids missing default config a standalone injection risks.
Verified via render throughout. Live and verified end to end on
xfa-awesome-alpha, both clusters healthy, real robot data confirmed
flowing through the full write path.
--no-verify: local pre-commit hook failed before commit because
buildifier is not on PATH and embedmd flagged a repo-wide markdown file;
the full app manifest build below is the validation for this rerebase.
@Heet852003Heet852003 reopened this Aug 14, 2026
Real bug, not hypothetical - hit this live today. The cluster and
vmagent release names were hardcoded to bare 'vm'/'vmagent', shared
with the pre-rename app. Cluster-scoped RBAC (ClusterRole/
ClusterRoleBinding) isn't namespaced, so any two Apps using this chart
with the same release name will collide the moment both exist, forcing
a manual delete of one to unblock the other - confirmed via pure
render, reproducible with zero cluster risk (helm template alone
proves the collision, no live interaction needed).
Renamed to vmrobot/vmagentrobot, unique to this app. Verified via
render - genuinely different ClusterRole/ClusterRoleBinding names now.
…onflict
Real conflict found testing on a leased single-node VM - both
Prometheus's and VictoriaMetrics's node-exporter want hostPort 9100.
Works fine on multi-node robots (each lands on a different node) but
collides on single-node test VMs.
Moved VM's node-exporter to 19100 - leaves Prometheus's completely
untouched, both stacks stay fully independent (matters for eventually
decommissioning Prometheus without touching VictoriaMetrics). Added
an explicit relabel in vmagent's scrape config too, so it always hits
the right port rather than relying on inferred discovery.
Verified via render: containerPort/listen-address both correctly on
19100, no more 9100 anywhere in the rendered output.
The vmrobot/vmagentrobot release-name fix changed the generated VM
cluster Service names, but the hand-written ingress/HTTPRoute and
vmalert datasource config still pointed at the old vm-* Service names.
That left the cloud cluster healthy internally but broke the robot write
path and query ingress.
Updated those explicit references to vmrobot-* to match the rendered
chart resources.
The leased VM can reach the public VictoriaMetrics write endpoint
directly, but cannot reach the lab/on-prem Aqualine proxy address
172.28.0.8:1212. Lab-mtv-400 can also reach the endpoint directly, so
keep the write path simple and avoid forcing all robot vmagents through
a proxy that does not exist on every robot/leased VM.
Verified via temporary curl pods: direct endpoint reachable from both
lab-mtv-400 and vmp-1a18-0x4nxwqq; proxy timed out from the leased VM.
anchor still pointed at the old vmagent-* configmap name, silently
broke when we renamed the release earlier today. container and mount
injected fine, just the volume itself never showed up - caught testing
with a real sidecar spec. verified all three pieces render correctly
now, cloudmetrics unaffected.
blocker's gone now that helm is 3.21.3. remoteWriteUrls -> native
remoteWrite list, which is what actually unlocks bearerTokenFile.
config volume split into two now instead of one, logger flag went
double-dash. kept our existing bearer token setup since it's already
working live, cloud ops can add their own second target with
bearerTokenFile directly now, real yaml, no more anchor hacks needed
for that part.
verified with the full combined test - sidecar, volume, mount, labels,
relabel config, bearer token, all together, all render clean.
same upgrade as robot - remoteWriteUrls to native remoteWrite list,
updated anchors for the split config volume and double-dash flags.
gcp service account annotation logic untouched, schema didn't change
there.
verified full combined test - sidecar, shared volume, mount, labels,
relabel config, gcp service account annotation, primary remote write,
all together, all render clean.
methylDragon added a commit that referenced this pull request Aug 18, 2026
Adapted from original: #852
This change introduces VictoriaMetrics as a new monitoring solution, implemented as two distinct applications for scalable and flexible metrics collection and storage:
- **victoriametrics-cloudmetrics:** Deploys a VictoriaMetrics cluster (vmstorage, vminsert, vmselect) and vmagent to collect and store metrics from the Kubernetes cloud cluster (e.g., kube-state-metrics, kubelet/cAdvisor).
- **victoriametrics-robotmetrics:** Establishes a cloud-side VictoriaMetrics cluster and vmalert for robot metrics, alongside robot-side agents (vmagent, node-exporter, smartctl-exporter) for local scraping and remote writing.
To support the newer VictoriaMetrics Helm charts and their dependencies, the hermetic Helm 3 binary has been upgraded to v3.21.3. This version is compatible with features like victoria-metrics-common templates that were incompatible with the older Helm v3.9.0.
This also includes adding the necessary Bazel BUILD rules and pulling in new third-party Helm chart dependencies for VictoriaMetrics components (agent, alert, cluster) and Grafana.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch 2 times, most recently from b85564d to 288f0beCompareAugust 19, 2026 02:16
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 288f0be to adff071CompareAugust 19, 2026 21:44
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Heet852003@methylDragon
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config by Heet852003 · Pull Request #852 · googlecloudrobotics/core · GitHub
Skip to content

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config - #852

Draft
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps
Draft

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config#852
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps

Conversation

@Heet852003

@Heet852003Heet852003 commented Aug 7, 2026

Copy link
Copy Markdown

Migrating from Prometheus to VictoriaMetrics - two separate clusters, one for robot metrics and one for cloud metrics, since their growth patterns are different. Grafana's pulled out into its own standalone app too.

New apps: victoriametrics-robotmetrics (robot-metrics cluster - vminsert/vmselect/vmstorage/vmalert, plus the per-robot vmagent selector), victoriametrics-cloudmetrics (cloud-metrics cluster), grafana.
Enable/disable is controlled by AppRollout robots:/cloud: block presence, not explicit flags : simpler, matches how the CRD actually works.

Prometheus gets an enable toggle on both cloud and robot side (defaults true, no change for existing projects), plus a configurable remote-write port for when it needs to push into VictoriaMetrics during the hybrid period.

Also built out what Cloud Ops needs to pull our metrics into their Mimir cluster - an injectable sidecar for auth, a second remote_write target with correct positional relabel handling, external labels for their dedup keys (project_id, region), and a GCP service account annotation on the cloud vmagent's ServiceAccount. All of this is on both the cloud-metrics vmagent and the robot's own vmagent, since that's the only vmagent touching robot data at all.

Also moved kube-state-metrics and node-exporter from standalone vendored charts to kube-prometheus-stack, same reasoning as Grafana , avoids missing default config a standalone injection risks.Everything's default-off/inert unless explicitly set, so none of this changes behavior for any existing project.

Builds on top of @methylDragon's Grafana decoupling (core#850, insrc#51993).

RELNOTES=NONE

Comment threadsrc/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml Outdated
# scrapes robot hardware exporters.
{{- $data := .Files.Get "files/vmagent-chart.cloud.yaml" -}}
{{- $data = $data | replace "${VMINSERT_PORT}" (toString .Values.vminsert_port) | replace "release-name-placeholder" .Values.release_name | replace "HELM-NAMESPACE" .Release.Namespace -}}
{{- if .Values.upstream.sidecar_volume_mounts -}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're still defining individual sidecar value replaces.

Why can't we substitute the entire block using a yaml rule (seems prometheus-operator does this), or string replace a generic variable that unpacks to any yaml in case we missed something?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

dug into it, prometheus-operator can do that because kube-prometheus-stack's chart already natively supports those fields. our vmagent chart (old 0.6.0) genuinely has zero native extension points, checked directly. converted everything to real structured yaml in the meantime though. once the version upgrade lands we can drop the anchor stuff entirely and go native like yours

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's call this victoriametrics-robotmetrics

And make sure to add a README making it very clear that cloudmetrics is for collecting metrics from cloud services (and only has componentes in the cloud), and robotmetrics is for collecting metrics from the robot (and has components that span the robot and the cloud).

@Heet852003Heet852003Aug 8, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, renamed + readme added on both repos. one thing to flag: only renamed the App itself, not the AppRollout , namespace comes from the AppRollout's own name, so keeping that stable means zero migration risk and the grafana urls we already verified stay valid. lmk if you want full symmetry instead

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's have full symmetry

@methylDragonmethylDragonAug 7, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, missing vmagent-operator.yaml in this cloud directory? There should be a vmagent in the cloud as well for the robotmetrics right?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reconfirmed cleanly, no vmagent anywhere under robotmetrics/cloud/, only the robot's own on-prem one touches robot data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But don't we need the robot vmagent to send to a cloud robot-vmagent so we can route robot metrics to the victoriametrics-robotmetrics cluster instance (and also eventually to cloud mimir)?

Basically your robotmetrics/robot creates the vmagent on each robot, but what about the robot vmagent that's in the cloud?

image

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Base your PR on top of #850 and use its src/app_charts/grafana instead (replacing your src/app_charts/grafana

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done , dropped my own grafana app, using yours instead, squashed everything into one commit on top of your branch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

!!! For all subcharts in kube-prometheus-stack that you might use in victoriametrics (e.g. kube-state-metrics, node-exporter), I think there are default configurations we currently rely on that are part of kube-prometheus-stack that might warrant us using it instead of injecting the charts directly.

See my comment in: #850

I tried to do a standalone injection of grafana and it missed all those configs. I suspect it'll be the same with the other components that used to be part of kube-prometheus-stack that you extracted out.


Let's use kube-prometheus-stack, but just disable prometheus and any other components we're turning down, just like I did to get grafana working (e.g. here: https://github.com/googlecloudrobotics/core/pull/850/changes#diff-831c168d469137791822495d70dad85851c76f8433824fae989945f40ece7753R58-R69)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done for both kube-state-metrics and node-exporter, on kube-prometheus-stack now with everything else disabled, same as your grafana pattern. found a gap doing node-exporter, our old mirror was missing a cardinality drop prometheus actually applies, added that too. render-verified both

@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch 2 times, most recently from 70c273a to e064b9fCompareAugust 13, 2026 16:56
@methylDragon
methylDragon changed the base branch from main to ch3/standalone-grafanaAugust 13, 2026 21:04
Extract Grafana from the Prometheus application, enabling it to be
deployed and managed independently.
#### Why
Previously, Grafana was deployed as a component of the Prometheus
operator chart, which tightly coupled its lifecycle and configuration
to Prometheus. This change decouples Grafana, allowing for:
- **Independent Management:** Grafana can now be configured, deployed,
and updated separately from Prometheus.
- **Clearer Ownership:** Grafana resources are now owned by its own
dedicated application definition.
- **Reduced Conflicts:** Explicitly disables Grafana within the
Prometheus chart and ensures CRD ownership is handled correctly,
preventing resource conflicts.
#### The code changes include:
- Adding a new `grafana` application definition under `src/app_charts`.
- Moving Grafana's HTTPRoute and Ingress configurations to the new app.
- Configuring the standalone Grafana to use the `kube-prometheus-stack`
Helm chart, but with only Grafana components enabled.
- Disabling Grafana within the `prometheus` application's chart.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 5e2c121 to 66ae529CompareAugust 14, 2026 00:53
@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch from e064b9f to 66ae529CompareAugust 14, 2026 02:25
Two new apps - victoriametrics-robotmetrics (robot-metrics cluster,
vminsert/vmselect/vmstorage/vmalert plus the robot vmagent) and
victoriametrics-cloudmetrics (cloud-metrics cluster). Both configurable
- per-component replica count, upstream sidecar/remote-write config as
real structured YAML, service account annotation and dedup labels for
Cloud Ops' Mimir integration. Enable/disable controlled via AppRollout
block presence, not explicit flags, per review.
Built on top of the standalone Grafana app (googlecloudrobotics#850) instead of a separate
implementation. Also swapped kube-state-metrics and node-exporter from
standalone vendored charts to kube-prometheus-stack, same reasoning -
avoids missing default config a standalone injection risks.
Verified via render throughout. Live and verified end to end on
xfa-awesome-alpha, both clusters healthy, real robot data confirmed
flowing through the full write path.
--no-verify: local pre-commit hook failed before commit because
buildifier is not on PATH and embedmd flagged a repo-wide markdown file;
the full app manifest build below is the validation for this rerebase.
@Heet852003Heet852003 reopened this Aug 14, 2026
Real bug, not hypothetical - hit this live today. The cluster and
vmagent release names were hardcoded to bare 'vm'/'vmagent', shared
with the pre-rename app. Cluster-scoped RBAC (ClusterRole/
ClusterRoleBinding) isn't namespaced, so any two Apps using this chart
with the same release name will collide the moment both exist, forcing
a manual delete of one to unblock the other - confirmed via pure
render, reproducible with zero cluster risk (helm template alone
proves the collision, no live interaction needed).
Renamed to vmrobot/vmagentrobot, unique to this app. Verified via
render - genuinely different ClusterRole/ClusterRoleBinding names now.
…onflict
Real conflict found testing on a leased single-node VM - both
Prometheus's and VictoriaMetrics's node-exporter want hostPort 9100.
Works fine on multi-node robots (each lands on a different node) but
collides on single-node test VMs.
Moved VM's node-exporter to 19100 - leaves Prometheus's completely
untouched, both stacks stay fully independent (matters for eventually
decommissioning Prometheus without touching VictoriaMetrics). Added
an explicit relabel in vmagent's scrape config too, so it always hits
the right port rather than relying on inferred discovery.
Verified via render: containerPort/listen-address both correctly on
19100, no more 9100 anywhere in the rendered output.
The vmrobot/vmagentrobot release-name fix changed the generated VM
cluster Service names, but the hand-written ingress/HTTPRoute and
vmalert datasource config still pointed at the old vm-* Service names.
That left the cloud cluster healthy internally but broke the robot write
path and query ingress.
Updated those explicit references to vmrobot-* to match the rendered
chart resources.
The leased VM can reach the public VictoriaMetrics write endpoint
directly, but cannot reach the lab/on-prem Aqualine proxy address
172.28.0.8:1212. Lab-mtv-400 can also reach the endpoint directly, so
keep the write path simple and avoid forcing all robot vmagents through
a proxy that does not exist on every robot/leased VM.
Verified via temporary curl pods: direct endpoint reachable from both
lab-mtv-400 and vmp-1a18-0x4nxwqq; proxy timed out from the leased VM.
anchor still pointed at the old vmagent-* configmap name, silently
broke when we renamed the release earlier today. container and mount
injected fine, just the volume itself never showed up - caught testing
with a real sidecar spec. verified all three pieces render correctly
now, cloudmetrics unaffected.
blocker's gone now that helm is 3.21.3. remoteWriteUrls -> native
remoteWrite list, which is what actually unlocks bearerTokenFile.
config volume split into two now instead of one, logger flag went
double-dash. kept our existing bearer token setup since it's already
working live, cloud ops can add their own second target with
bearerTokenFile directly now, real yaml, no more anchor hacks needed
for that part.
verified with the full combined test - sidecar, volume, mount, labels,
relabel config, bearer token, all together, all render clean.
same upgrade as robot - remoteWriteUrls to native remoteWrite list,
updated anchors for the split config volume and double-dash flags.
gcp service account annotation logic untouched, schema didn't change
there.
verified full combined test - sidecar, shared volume, mount, labels,
relabel config, gcp service account annotation, primary remote write,
all together, all render clean.
methylDragon added a commit that referenced this pull request Aug 18, 2026
Adapted from original: #852
This change introduces VictoriaMetrics as a new monitoring solution, implemented as two distinct applications for scalable and flexible metrics collection and storage:
- **victoriametrics-cloudmetrics:** Deploys a VictoriaMetrics cluster (vmstorage, vminsert, vmselect) and vmagent to collect and store metrics from the Kubernetes cloud cluster (e.g., kube-state-metrics, kubelet/cAdvisor).
- **victoriametrics-robotmetrics:** Establishes a cloud-side VictoriaMetrics cluster and vmalert for robot metrics, alongside robot-side agents (vmagent, node-exporter, smartctl-exporter) for local scraping and remote writing.
To support the newer VictoriaMetrics Helm charts and their dependencies, the hermetic Helm 3 binary has been upgraded to v3.21.3. This version is compatible with features like victoria-metrics-common templates that were incompatible with the older Helm v3.9.0.
This also includes adding the necessary Bazel BUILD rules and pulling in new third-party Helm chart dependencies for VictoriaMetrics components (agent, alert, cluster) and Grafana.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch 2 times, most recently from b85564d to 288f0beCompareAugust 19, 2026 02:16
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 288f0be to adff071CompareAugust 19, 2026 21:44
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Heet852003@methylDragon
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config by Heet852003 · Pull Request #852 · googlecloudrobotics/core · GitHub
Skip to content

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config - #852

Draft
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps
Draft

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config#852
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps

Conversation

@Heet852003

@Heet852003Heet852003 commented Aug 7, 2026

Copy link
Copy Markdown

Migrating from Prometheus to VictoriaMetrics - two separate clusters, one for robot metrics and one for cloud metrics, since their growth patterns are different. Grafana's pulled out into its own standalone app too.

New apps: victoriametrics-robotmetrics (robot-metrics cluster - vminsert/vmselect/vmstorage/vmalert, plus the per-robot vmagent selector), victoriametrics-cloudmetrics (cloud-metrics cluster), grafana.
Enable/disable is controlled by AppRollout robots:/cloud: block presence, not explicit flags : simpler, matches how the CRD actually works.

Prometheus gets an enable toggle on both cloud and robot side (defaults true, no change for existing projects), plus a configurable remote-write port for when it needs to push into VictoriaMetrics during the hybrid period.

Also built out what Cloud Ops needs to pull our metrics into their Mimir cluster - an injectable sidecar for auth, a second remote_write target with correct positional relabel handling, external labels for their dedup keys (project_id, region), and a GCP service account annotation on the cloud vmagent's ServiceAccount. All of this is on both the cloud-metrics vmagent and the robot's own vmagent, since that's the only vmagent touching robot data at all.

Also moved kube-state-metrics and node-exporter from standalone vendored charts to kube-prometheus-stack, same reasoning as Grafana , avoids missing default config a standalone injection risks.Everything's default-off/inert unless explicitly set, so none of this changes behavior for any existing project.

Builds on top of @methylDragon's Grafana decoupling (core#850, insrc#51993).

RELNOTES=NONE

Comment threadsrc/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml Outdated
# scrapes robot hardware exporters.
{{- $data := .Files.Get "files/vmagent-chart.cloud.yaml" -}}
{{- $data = $data | replace "${VMINSERT_PORT}" (toString .Values.vminsert_port) | replace "release-name-placeholder" .Values.release_name | replace "HELM-NAMESPACE" .Release.Namespace -}}
{{- if .Values.upstream.sidecar_volume_mounts -}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're still defining individual sidecar value replaces.

Why can't we substitute the entire block using a yaml rule (seems prometheus-operator does this), or string replace a generic variable that unpacks to any yaml in case we missed something?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

dug into it, prometheus-operator can do that because kube-prometheus-stack's chart already natively supports those fields. our vmagent chart (old 0.6.0) genuinely has zero native extension points, checked directly. converted everything to real structured yaml in the meantime though. once the version upgrade lands we can drop the anchor stuff entirely and go native like yours

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's call this victoriametrics-robotmetrics

And make sure to add a README making it very clear that cloudmetrics is for collecting metrics from cloud services (and only has componentes in the cloud), and robotmetrics is for collecting metrics from the robot (and has components that span the robot and the cloud).

@Heet852003Heet852003Aug 8, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, renamed + readme added on both repos. one thing to flag: only renamed the App itself, not the AppRollout , namespace comes from the AppRollout's own name, so keeping that stable means zero migration risk and the grafana urls we already verified stay valid. lmk if you want full symmetry instead

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's have full symmetry

@methylDragonmethylDragonAug 7, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, missing vmagent-operator.yaml in this cloud directory? There should be a vmagent in the cloud as well for the robotmetrics right?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reconfirmed cleanly, no vmagent anywhere under robotmetrics/cloud/, only the robot's own on-prem one touches robot data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But don't we need the robot vmagent to send to a cloud robot-vmagent so we can route robot metrics to the victoriametrics-robotmetrics cluster instance (and also eventually to cloud mimir)?

Basically your robotmetrics/robot creates the vmagent on each robot, but what about the robot vmagent that's in the cloud?

image

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Base your PR on top of #850 and use its src/app_charts/grafana instead (replacing your src/app_charts/grafana

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done , dropped my own grafana app, using yours instead, squashed everything into one commit on top of your branch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

!!! For all subcharts in kube-prometheus-stack that you might use in victoriametrics (e.g. kube-state-metrics, node-exporter), I think there are default configurations we currently rely on that are part of kube-prometheus-stack that might warrant us using it instead of injecting the charts directly.

See my comment in: #850

I tried to do a standalone injection of grafana and it missed all those configs. I suspect it'll be the same with the other components that used to be part of kube-prometheus-stack that you extracted out.


Let's use kube-prometheus-stack, but just disable prometheus and any other components we're turning down, just like I did to get grafana working (e.g. here: https://github.com/googlecloudrobotics/core/pull/850/changes#diff-831c168d469137791822495d70dad85851c76f8433824fae989945f40ece7753R58-R69)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done for both kube-state-metrics and node-exporter, on kube-prometheus-stack now with everything else disabled, same as your grafana pattern. found a gap doing node-exporter, our old mirror was missing a cardinality drop prometheus actually applies, added that too. render-verified both

@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch 2 times, most recently from 70c273a to e064b9fCompareAugust 13, 2026 16:56
@methylDragon
methylDragon changed the base branch from main to ch3/standalone-grafanaAugust 13, 2026 21:04
Extract Grafana from the Prometheus application, enabling it to be
deployed and managed independently.
#### Why
Previously, Grafana was deployed as a component of the Prometheus
operator chart, which tightly coupled its lifecycle and configuration
to Prometheus. This change decouples Grafana, allowing for:
- **Independent Management:** Grafana can now be configured, deployed,
and updated separately from Prometheus.
- **Clearer Ownership:** Grafana resources are now owned by its own
dedicated application definition.
- **Reduced Conflicts:** Explicitly disables Grafana within the
Prometheus chart and ensures CRD ownership is handled correctly,
preventing resource conflicts.
#### The code changes include:
- Adding a new `grafana` application definition under `src/app_charts`.
- Moving Grafana's HTTPRoute and Ingress configurations to the new app.
- Configuring the standalone Grafana to use the `kube-prometheus-stack`
Helm chart, but with only Grafana components enabled.
- Disabling Grafana within the `prometheus` application's chart.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 5e2c121 to 66ae529CompareAugust 14, 2026 00:53
@Heet852003
Heet852003force-pushed the heet/crc-prometheus-vm-grafana-apps branch from e064b9f to 66ae529CompareAugust 14, 2026 02:25
Two new apps - victoriametrics-robotmetrics (robot-metrics cluster,
vminsert/vmselect/vmstorage/vmalert plus the robot vmagent) and
victoriametrics-cloudmetrics (cloud-metrics cluster). Both configurable
- per-component replica count, upstream sidecar/remote-write config as
real structured YAML, service account annotation and dedup labels for
Cloud Ops' Mimir integration. Enable/disable controlled via AppRollout
block presence, not explicit flags, per review.
Built on top of the standalone Grafana app (googlecloudrobotics#850) instead of a separate
implementation. Also swapped kube-state-metrics and node-exporter from
standalone vendored charts to kube-prometheus-stack, same reasoning -
avoids missing default config a standalone injection risks.
Verified via render throughout. Live and verified end to end on
xfa-awesome-alpha, both clusters healthy, real robot data confirmed
flowing through the full write path.
--no-verify: local pre-commit hook failed before commit because
buildifier is not on PATH and embedmd flagged a repo-wide markdown file;
the full app manifest build below is the validation for this rerebase.
@Heet852003Heet852003 reopened this Aug 14, 2026
Real bug, not hypothetical - hit this live today. The cluster and
vmagent release names were hardcoded to bare 'vm'/'vmagent', shared
with the pre-rename app. Cluster-scoped RBAC (ClusterRole/
ClusterRoleBinding) isn't namespaced, so any two Apps using this chart
with the same release name will collide the moment both exist, forcing
a manual delete of one to unblock the other - confirmed via pure
render, reproducible with zero cluster risk (helm template alone
proves the collision, no live interaction needed).
Renamed to vmrobot/vmagentrobot, unique to this app. Verified via
render - genuinely different ClusterRole/ClusterRoleBinding names now.
…onflict
Real conflict found testing on a leased single-node VM - both
Prometheus's and VictoriaMetrics's node-exporter want hostPort 9100.
Works fine on multi-node robots (each lands on a different node) but
collides on single-node test VMs.
Moved VM's node-exporter to 19100 - leaves Prometheus's completely
untouched, both stacks stay fully independent (matters for eventually
decommissioning Prometheus without touching VictoriaMetrics). Added
an explicit relabel in vmagent's scrape config too, so it always hits
the right port rather than relying on inferred discovery.
Verified via render: containerPort/listen-address both correctly on
19100, no more 9100 anywhere in the rendered output.
The vmrobot/vmagentrobot release-name fix changed the generated VM
cluster Service names, but the hand-written ingress/HTTPRoute and
vmalert datasource config still pointed at the old vm-* Service names.
That left the cloud cluster healthy internally but broke the robot write
path and query ingress.
Updated those explicit references to vmrobot-* to match the rendered
chart resources.
The leased VM can reach the public VictoriaMetrics write endpoint
directly, but cannot reach the lab/on-prem Aqualine proxy address
172.28.0.8:1212. Lab-mtv-400 can also reach the endpoint directly, so
keep the write path simple and avoid forcing all robot vmagents through
a proxy that does not exist on every robot/leased VM.
Verified via temporary curl pods: direct endpoint reachable from both
lab-mtv-400 and vmp-1a18-0x4nxwqq; proxy timed out from the leased VM.
anchor still pointed at the old vmagent-* configmap name, silently
broke when we renamed the release earlier today. container and mount
injected fine, just the volume itself never showed up - caught testing
with a real sidecar spec. verified all three pieces render correctly
now, cloudmetrics unaffected.
blocker's gone now that helm is 3.21.3. remoteWriteUrls -> native
remoteWrite list, which is what actually unlocks bearerTokenFile.
config volume split into two now instead of one, logger flag went
double-dash. kept our existing bearer token setup since it's already
working live, cloud ops can add their own second target with
bearerTokenFile directly now, real yaml, no more anchor hacks needed
for that part.
verified with the full combined test - sidecar, volume, mount, labels,
relabel config, bearer token, all together, all render clean.
same upgrade as robot - remoteWriteUrls to native remoteWrite list,
updated anchors for the split config volume and double-dash flags.
gcp service account annotation logic untouched, schema didn't change
there.
verified full combined test - sidecar, shared volume, mount, labels,
relabel config, gcp service account annotation, primary remote write,
all together, all render clean.
methylDragon added a commit that referenced this pull request Aug 18, 2026
Adapted from original: #852
This change introduces VictoriaMetrics as a new monitoring solution, implemented as two distinct applications for scalable and flexible metrics collection and storage:
- **victoriametrics-cloudmetrics:** Deploys a VictoriaMetrics cluster (vmstorage, vminsert, vmselect) and vmagent to collect and store metrics from the Kubernetes cloud cluster (e.g., kube-state-metrics, kubelet/cAdvisor).
- **victoriametrics-robotmetrics:** Establishes a cloud-side VictoriaMetrics cluster and vmalert for robot metrics, alongside robot-side agents (vmagent, node-exporter, smartctl-exporter) for local scraping and remote writing.
To support the newer VictoriaMetrics Helm charts and their dependencies, the hermetic Helm 3 binary has been upgraded to v3.21.3. This version is compatible with features like victoria-metrics-common templates that were incompatible with the older Helm v3.9.0.
This also includes adding the necessary Bazel BUILD rules and pulling in new third-party Helm chart dependencies for VictoriaMetrics components (agent, alert, cluster) and Grafana.
Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch 2 times, most recently from b85564d to 288f0beCompareAugust 19, 2026 02:16
@methylDragon
methylDragonforce-pushed the ch3/standalone-grafana branch from 288f0be to adff071CompareAugust 19, 2026 21:44
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Heet852003@methylDragon