Conversation
There was a problem hiding this comment.
🟡 Changes recommended
The new tombstone and iterator code paths need additional defensive validation and invariant enforcement to avoid incorrect behavior or potential out-of-bounds access on manufactured/invalid IDs.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
This PR adds a “tombstone” mechanism to ts::Metrics so a metric can be withdrawn from publication (hidden from iteration/enum-based consumers) while remaining resolvable by exact name / id and keeping its backing atomic/value.
Changes:
- Extend
ts::Metrics::Storagewith a per-slot flag array and publictombstone()/tombstoned()APIs. - Update
ts::Metricsiteration semantics to skip tombstoned slots and to use a snapshot bound captured at iterator construction. - Add unit tests covering tombstoning behavior across both
ts::Metricsand records lookup, plus documentation updates.
File summaries
| File | Description |
|---|---|
src/tsutil/Metrics.cc |
Implements tombstone flagging and iterator behavior changes (snapshot bound + skip). |
include/tsutil/Metrics.h |
Exposes the tombstone API, adds per-slot flag storage, and updates iterator semantics/contracts. |
src/tsutil/unit_tests/test_Metrics.cc |
Adds coverage for tombstone behavior, iteration skipping, resurrection, and edge cases. |
src/records/unit_tests/test_RecHiddenMetricLookup.cc |
Verifies record lookup behavior with tombstoned metrics (enumeration vs exact lookup). |
doc/developer-guide/internal-libraries/Metrics.en.rst |
Documents the tombstone feature and its interaction with find()/iteration. |
Review details
Suppressed comments (2)
src/tsutil/Metrics.cc:288
Storage::tombstoned()indexes the per-slot flag array withoffsetwithout validating thatoffset < MAX_SIZE(or that the slot is allocated in the current blob). A manufactured/invalidIdTypewith a large offset can trigger out-of-bounds access; it should safely return false for non-allocated/non-sensical IDs.
Metrics::Storage::tombstoned(Metrics::IdType id) const
{
auto [blob_ix, offset] = _splitID(id);
Metrics::NamesAndAtomics *blob = _blobs[blob_ix].get();
if (!blob) {
return false;
}
return (std::get<2>(*blob)[offset].load(MEMORY_ORDER) & TOMBSTONE) != 0;
}
src/tsutil/Metrics.cc:296
- The positional iterator ctor
iterator(const Metrics&, IdType)does not callskip_tombstoned(). That allows external callers to construct an iterator that points at a tombstoned slot (contradicting the intended "iteration never visits marked slots" invariant) and reintroduces the non-terminating range-walk risk if such an iterator is used as a bound.
Metrics::iterator::iterator(const Metrics &m, IdType pos) : _metrics(m), _it(pos), _bound(m._storage->current_id()) {}
- Files reviewed: 5/5 changed files
- Comments generated: 1
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
Thanks — all three findings were real, including the two that were filed as suppressed comments. Fixed in 9ed3ee8.
bool
allocated(IdType id) const
{
auto [blob, entry] = _splitID(id);
if (id < 0 || entry >= MAX_SIZE || !_blobs[blob]) {
return false;
}
return blob < _cur_blob || (blob == _cur_blob && entry < _cur_off);
}This is deliberately stricter than the existing I left The positional iterator constructor. You are right that it let a caller rest an iterator on a tombstoned slot, which reintroduces the non-terminating range-walk. Rather than only skipping, the three constructors are now private with New tests covering each case: an offset past the end of a full blob, the next free slot, and a blob index that was never allocated. |
|
Follow-up on a review question: That is in fact safe, but only by coincidence. b4fee22 makes it explicit: static_assert(MAX_BLOBS == METRIC_TYPE_MASK + 1, "a masked blob index must always be a valid _blobs index");Verified it fires: dropping Also added a test for the largest possible id, which exercises the other half — the offset is not masked to the blob size, so it needs the explicit Worth noting for a possible follow-up, out of scope here: |
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated 2 comments.
|
|
||
| return blob < _cur_blob || (blob == _cur_blob && entry < _cur_off); | ||
| } | ||
| }; | ||
|
|
||
| Metrics(std::shared_ptr<Storage> &str) : _storage(str) {} | ||
|
|
||
| std::shared_ptr<Storage> _storage; | ||
|
|
||
| public: | ||
| // These are sort of factory classes, using the Metrics singleton for all storage etc. |
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated 3 comments.
There was a problem hiding this comment.
🔵 Needs a closer look
It changes core metrics storage/iteration semantics used broadly across ATS, and needs human validation with full CI results and concurrency/compatibility scrutiny.
Review details
- Files reviewed: 5/5 changed files
- Comments generated: 0 new
- Review effort level: Lite
|
Went through all six. Two were already addressed, three are the same point and are now fixed, and one is incorrect — details below. Fixed in ef97aa6
Not a defect: the negative-offset concern
The Already addressed The The PR description mismatch was real when filed; the description now documents |
There was a problem hiding this comment.
🟡 Changes recommended
The new per-slot atomic flag storage needs explicit initialization to avoid nondeterministic “unlisted” state, and one unit-test section can dereference begin() after making the store empty (undefined behavior).
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
src/tsutil/unit_tests/test_Metrics.cc:740
- This section dereferences *m.begin() after unlisting id 0; if this TEST_CASE path runs before any other metric is created, begin()==end() and dereferencing is undefined. Create a guaranteed-listed metric first so begin() is always safe to dereference.
- Files reviewed: 5/5 changed files
- Comments generated: 1
- Review effort level: Lite
| if (it != _lookups.end()) { | ||
| // Re-creating a name is how an unlisted metric is relisted: same slot, same atomic, and | ||
| // whatever value it accumulated while it was out of the listing. A name in _lookups always | ||
| // names an allocated slot, so this cannot fail. | ||
| set_listed(it->second, true); | ||
|
|
That was true through C++17, but P0883 changed It also does not depend on that change. I would rather not argue that from the standard, so 02d30f9 asserts it instead. The blob growth test already creates Confirmed the assertion is not vacuous. Storing That matters more than the standard argument, because reading uninitialized heap frequently does return zero — fresh pages are zero-filled — so this class of bug hides well and a test that only samples a metric or two would not catch it. I did not add an explicit initialization loop. It would be dead work on every blob, and the real risk is not today's behavior but a future change to something like |
There was a problem hiding this comment.
🟡 Changes recommended
Metrics::iterator can incorrectly treat listed GAUGE metrics as “end” due to type-bit contamination in find() positional iterators when using the new _bound numeric comparison.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 5/5 changed files
- Comments generated: 1
- Review effort level: Lite
02d30f9 to
7b19905
Compare
|
This one is a real bug, and mine. Fixed in 7b19905.
It is a regression from the bound comparison I introduced with unlisting — before that, Fixed as suggested, by keeping only the blob and offset: auto [blob, offset] = _metrics._splitID(pos);
_it = _makeId(blob, offset, MetricType::COUNTER);Dereferencing is unaffected: Test added first and watched fail on Separately, I have dropped the commit I added earlier about |
There was a problem hiding this comment.
🟡 Changes recommended
The new allocated()-based checks introduce unsynchronized reads of shared storage state that can be exercised by the new public APIs, creating a C++ data race risk under concurrent metric registration.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 5/5 changed files
- Comments generated: 1
- Review effort level: Lite
|
Correct as stated, but deliberately out of scope here.
So this is not a race introduced by unlisting; it is the existing synchronization model of The synchronization of this class is being addressed directly in #13583, "Metrics: close the id lookup race and bounds gaps left by the lock revert". That is the right place for it — it is a property of the whole store, not of this feature, and fixing it in two PRs at once would just produce conflicts. Whichever of the two lands second should extend the fix to cover the other's accessors: if #13583 goes first, |
Each iterator captures its own bound, and exhaustion was judged against that. A subrange whose stop iterator was made later held a larger bound, so the walk could pass its own bound and go on comparing unequal to a stop that was still live, with operator++ unable to make progress. Two find() calls with a metric created between them was enough. Exhaustion between two positional iterators is now judged against the earlier of the two bounds, so such a subrange ends at the earlier snapshot. The sentinel keeps its own answer, since its bound means nothing.
It asserted the store had at least one listed metric left, which depends on what other sections put there. A listed metric of its own says the same thing without that coupling.
565db22 to
1f0f4ca
Compare
|
Five comments, one real bug. Taking them in order of consequence. The subrange across snapshots — real, fixed in e776b65 This is a genuine defect and reachable through the public API. Two auto start = m.find("a"); // captures bound B1
Metrics::Counter::create("z"); // store grows
auto stop = m.find("z"); // bound B2 > B1, and its position is >= B1
for (auto it = start; it != stop; ++it) { ... } // never terminates
Fixed by judging exhaustion between two positional iterators against the earlier of the two bounds, so such a subrange ends at the earlier snapshot. The sentinel keeps its own answer, since its bound is meaningless — folding it into the minimum would make every other iterator compare exhausted immediately. My earlier subrange test missed this because it created both endpoints after all the metrics, so both held the same bound. The Filed twice, and right both times: that assertion did depend on the shared store holding at least one listed metric. In practice slot 0 is always listed so it could not actually fail, but the coupling is real and pointless. The section now creates a listed metric of its own and asserts it is observed. One correction to the suggestion: an anchor placed before the unlisted tail proves the loop ran, not that iteration reached the tail. The "none of the tail names appear" assertion is what covers the tail. The comment says only the former. Structured bindings and The concern is that binding and is clean. The pattern is also already in the tree — two occurrences in this same file before this PR, three in
Correct that Both commits build and pass |
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated 2 comments.
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated 1 comment.
Suppressed comments (1)
src/tsutil/Metrics.cc:1
iterator(const Metrics&, IdType)normalizes_ittoMetricType::COUNTER, butStorage::type()derives the type from the id bits (_extractType(id)). If the iterator id is coerced to COUNTER, iteration andfind()dereference can report the wrong metric type (e.g., GAUGE/COUNTER mismatches) and potentially break any logic that depends on the encoded type. A robust fix is to keep_itas the canonical id (with its real type bits) and compare exhaustion/ordering using a separate positional value (e.g., blob+offset linear index), or to fetch the type for dereference from the slot’s stored id rather than from_it.
/** @file
Exhaustion is a property of an iterator's own snapshot bound, so two taken at different times can compare equal to each other while disagreeing about end. That is not a total equivalence relation, which makes these unfit for a generic algorithm; only same snapshot comparisons, and comparison against end, are meaningful.
|
Two to answer, since the Iterator equality is not an equivalence relation — correct, documented in 03452ff The counterexample is real and I can be precise about it. With bounds The cause is structural rather than a slip in the comparison: exhaustion is a property of each iterator's own snapshot, so any iterator-versus-end test is per-iterator, and no comparison rule over a single type can paper over that. I looked at the sentinel design you suggest. It does fix it properly — iterator-versus-iterator becomes pure position equality, iterator-versus-sentinel carries the end test, and the two are different relations so nothing is required to hold across them. The cost is that We decided against it for now and documented the limitation instead. These are Worth adding: the min-bound rule this replaced a hang, not correct behaviour. Before it, that same subrange did not terminate at all. Many SECTIONs mutating the singleton store — declining The remedy does not do what it is meant to here. Splitting into separate TEST_CASE("Metrics unlisting", "[libtsapi][Metrics]")
{
auto &m = Metrics::instance();
SECTION(...)The body is a reference binding. Every registration happens inside a On volume: the largest fill in this test case is 8 metrics, and it adds a few dozen in total against a store that holds 8M. The pattern of many sections over The one place this was a real problem was an assertion that depended on the store holding at least one listed metric from elsewhere, which is fixed in 1f0f4ca. |
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.
The previous note disclaimed the equivalence relation while the type still declared input_iterator_tag, which advertises what it then denied. A snapshot is the sequence: iterators from different ones are no more comparable than iterators into different containers, so mixing them is unspecified rather than broken, and within one snapshot equality is the relation an input iterator requires.
|
Both of these arrived twice; answering once each.
The objection lands on the previous wording rather than on the code: declaring The framing that fits what the code does: a snapshot is the sequence. The min-bound rule then has a narrower job than the old note implied: it makes the out-of-domain case terminate instead of hang. Before it, that comparison did not terminate at all, which is a worse kind of unspecified. On the two suggested remedies: the sentinel design does fix the relation properly and I priced it out — Private iterator constructors are a source break — correct, and now documented Accurate: I am keeping it private. Nothing in tree constructs an iterator directly, and there is no sensible reason to — a caller holding an id wants What was missing was disclosure, so the PR description now has an API change section recording it, for the 11.0.0 release notes alongside the |
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Pull request overview
Copilot reviewed 5 out of 5 changed files in this pull request and generated 2 comments.
|
Follow-up #13666 opens the consumer of |
bneradt
left a comment
There was a problem hiding this comment.
Two correctness concerns with changes to listing state during iterator use. I validated both using the iterator/find implementation from this revision in an isolated C++20 harness with a mock store; I did not run the full ATS test suite. The retained metric storage is an intentional, documented limitation and is not a finding.
| if (id == NOT_FOUND || !listed(id)) { | ||
| return end(); | ||
| } else { | ||
| return iterator(*this, id); |
There was a problem hiding this comment.
[P2] Keep find() from returning a different metric after concurrent unlisting
There is a check/use gap between listed(id) above and the positional constructor's skip_unlisted(). With adjacent listed metrics a and b, find("a") can observe listed(a) == true, another thread can unlist a, and the constructor then advances to b. The returned iterator is non-end but dereferences to a metric whose name does not match the query. I reproduced that interleaving by triggering the unlist when the constructor acquires its bound. Please make the find path retain the requested position or return end if that position is skipped, rather than accepting the next listed metric, and add a regression test for this interleaving.
There was a problem hiding this comment.
Confirmed the check/use gap: listed(id) here and skip_unlisted() in the positional constructor are two separate decisions, so an unlist in between makes the returned iterator dereference to the next listed metric instead of the one named.
Rather than make the find path retain the requested position, the positional constructor and find() are both gone as of f340f14. Enumeration is now Metrics::for_each(func) and there is no public iterator at all, so there is no position for a caller to name and nothing to race against. lookup() is the way to reach a single metric by name, which is what RecLookupRecord, LogAccess and TSStatFindName already used.
Worth noting find() predated this PR and was safe there -- the constructor did not skip, and equality was a plain position compare. This PR is what made it unsafe, which is why removing it here rather than in a separate change seemed right. It has no callers in src, include, plugins, example or tools; the only ones were the tests added by this PR. Do you know of out-of-tree consumers? Metrics.h is in TSUTIL_PUBLIC_HEADERS so this is a source-breaking removal, and I would rather hear about users now than after 11.0.0.
| return at_end() == o.at_end(); | ||
| } | ||
|
|
||
| auto const bound = _bound < o._bound ? _bound : o._bound; |
There was a problem hiding this comment.
[P2] Account for a saved subrange endpoint becoming unlisted
The shared-bound rule does not cover listing changes within the same allocation snapshot. Create a, b, and c, obtain start = m.find("a") and stop = m.find("b"), then call m.unlist("b") without creating any metrics. A walk for (auto it = start; it != stop; ++it) skips b, visits c outside the intended range, and passes its end without ever comparing equal to stop: the saved stop remains below the shared bound while the walking iterator is exhausted. This also arises when another thread unlists the endpoint during traversal. The current subrange test only unlists interior slots before obtaining the iterators. Please either support this endpoint transition or explicitly define and enforce the iterator invalidation/synchronization requirement for unlisting, with a regression test; snapshotting only the allocation bound does not preserve the endpoint.
There was a problem hiding this comment.
Reproduced, and it goes one step further than described. With a, b, c listed, start = find("a"), stop = find("b"), then unlist("b"): the walk visits a, skips the unlisted b and lands on c, compares live-vs-live so it visits c outside the range, then exhausts. At that point at_end() is true for the walker and false for the saved stop, so a && b is false and != stays true while ++ cannot make progress -- it spins, dereferencing _it == _bound. Storage::lookup guards with offset > _cur_off, so one-past-end slips past the redirect and yields an empty name rather than bad_id. So: out-of-range visit followed by a non-terminating loop.
Your reading of the coverage was right too -- the subrange test only unlisted interior slots, and only before obtaining the iterators. Neither test unlisted an endpoint after capturing it.
I went with the second of your two options, but by construction rather than by contract: as of f340f14 there is no way to form a subrange. Metrics::for_each(func) is the only enumeration, always the whole store, so there is no cursor to save across a listing change and no invalidation rule to define or enforce. _bound, the shared-bound comparison and the equality rules are all gone with the iterator. The two subrange tests are deleted rather than extended, since what they covered can no longer be expressed.
Net effect on the PR is 170 insertions against 344 deletions.
A public iterator lets a caller name a position, and a position stops meaning anything once iteration skips unlisted slots. An iterator held at a slot that is later unlisted becomes a range bound the walk steps straight over and never reaches, and find() could hand back the next listed metric rather than the one asked for. Supporting either would mean defining iterator invalidation for listing changes, to keep a surface with no callers: every consumer walks the whole store, and find() had none at all. for_each is the whole store or nothing. With no cursor to outlive the walk, the equality rules, the snapshot bound comparison and find() go away along with the defects they carried. lookup() remains the way to reach a single metric by name.
There was a problem hiding this comment.
🟡 Changes recommended
Add the direct <utility> include and document the public API removals and migration path.
Get a fresh assessment by requesting another Copilot review.
Review details
- Files reviewed: 6/6 changed files
- Comments generated: 2
- Review effort level: Lite
| } else { | ||
| return iterator(*this, id); | ||
| } | ||
| _storage->for_each(std::forward<F>(func)); |
| /** Visit every listed metric. | ||
| * | ||
| * @a func is called as <tt>func(std::string_view name, MetricType type, int64_t value)</tt> for | ||
| * each listed metric, in creation order. Unlisted metrics are skipped, @see unlist. | ||
| * | ||
| * The set walked is fixed when the call begins: a metric created while it runs is not visited. | ||
| * Enumeration is deliberately the whole store and nothing less. There is no cursor to hold, so | ||
| * nothing can outlive the walk or name a slot the walk would not visit, and @a func may not | ||
| * create a metric, which would be an attempt to grow the store from inside a pass over it. |
Problem
ts::Metrics::Storagehas no removal path.create()allocates a slot and a name and nothing ever undoes either, so a metric name lives for the life of the process.Any code that decides whether to publish a name based on a runtime changeable input therefore makes a permanent commitment the first time it publishes. The decision is latched at first creation and can never be revisited.
The case that surfaced this is the per upstream server connection metrics from #13506.
proxy.config.http.per_server.connection.metric_aggregateisRECU_DYNAMICand overridable, and at value2the per group<fqdn>.<ip>:<port>metrics are supposed to stay hidden while only the per hostname aggregates are published. On a box that ran for a while at0before being switched to2, both shapes are present intraffic_ctl metric match per_server, and no reload can remove the first set. The config change took effect correctly for everything created after it; the names created before it simply cannot be withdrawn.What this does
Lets a metric be taken out of the store's listing.
An unlisted metric:
traffic_ctl metric match, the JSONRPC record lookup andstats_over_httpwith no change in any of those consumers;lookup(), soRecLookupRecord,LogAccessfield resolution andTSStatFindNamekeep working;Derivedaggregate sourcing from it is unaffected;create()on the same name, returning the same id with its accumulated value intact.An unlisted phone number is the analogy: not in the directory, but it still rings if you know it. This is a publication policy, not a lifetime — any
IdTypeorAtomicType *a caller already holds stays valid across an unlist and relist.This PR adds the mechanism only. Nothing in the tree calls it, so every existing metric enumerates exactly as before. The
ConnectionTrackerfix is a follow up.Why enumeration had to be refactored
Unlisting means enumeration skips slots. Once it skips, a position in the store is no longer a meaningful place in the sequence — and a public iterator is nothing but a position a caller can hold onto.
Earlier revisions of this PR kept the iterator and tried to make positions safe under skipping. That does not work, and review found two ways it fails, both reproduced by @bneradt:
find()checkedlisted(id)and then the positional constructor skipped forward from it. An unlist between the two makes the returned iterator dereference to the next listed metric rather than the one named.++cannot make progress, and it dereferences one past the end of the allocated slots.The two fixes pulled against each other. Copilot flagged that the positional constructor not skipping let a caller name an unlisted slot, which enumeration must never visit; adding the skip to satisfy that is exactly what created the
find()defect above. Skipping enumeration and positional bounds are in tension and you cannot have both.So enumeration is now the whole store or nothing:
With no cursor to hand out, there is nothing to outlive a walk, nothing to name a slot the walk would skip, and no invalidation rule to define or enforce.
Metrics::iterator,begin(),end(),find(), the snapshot bound, the shared bound equality rule and the id type bit normalization all go away with the defects they carried.lookup()is the way to reach a single metric by name.Nothing wanted the removed surface. There are three enumeration sites in the tree, all in
RecCore.cc, and all are full passes that filter per element.find()had no callers at all outside its own tests — not insrc,include,plugins,exampleortools.Net effect of the refactor is 170 insertions against 344 deletions.
On the naming
This started out called
tombstone, which was wrong twice over. A tombstone elsewhere is a record that something was deleted, and this codebase already uses it that way —CacheShmtombstones a slot to mark it dead and reusable. Nothing is deleted here.hide/publishwould read best in isolation but both words are already load bearing for a different mechanism in this same class: the two separate stores,hidden_instance()andcreateHiddenPtr()versus the published store. An unlisted metric in the published store would have been "published but not published".listedcollides with neither, and says the useful part out loud.Implementation notes
Storage. A parallel
FlagStoragearray in the blob, rather than a member ofNameAndId: anstd::atomicmember would make that tuple neither copyable nor movable, and the slot is written with a tuple assignment. Blobs are already built withmake_unique, which value initializes, so flags start zero with no change toaddBlob(). Reads are lock free at relaxed ordering, matching the rest of the class. Cost is 1 KiB per 1024 slots against a blob that is already about 48 KiB. The singleUNLISTEDbit is set and cleared withfetch_or/fetch_andrather than a whole word store, so a flag added later is not clobbered.The walk.
Storage::for_eachreadsnext_free_id()once. That load is the sequence: acquiring it acquires every slot below it, which is what lets the walk read names and values without the mutex — a slot's name is written before the release store that publishes it and never changes afterwards. It then scans blob by blob up to that bound, skips any slot withUNLISTEDset, and takes each metric's type from the slot's own stored id rather than from its position, so a gauge is always reported as a gauge. A metric created while a walk runs is simply below no bound it agreed to visit.Metrics::BAD_ID_NAMEnames the reserved slot 0, which both theStorageconstructor andRecCorenow use. The hidden store enumeration used to skip that slot positionally, withbegin(); ++it;; it now skips by name, which is what it was actually expressing.Id validation.
Storage::_is_allocated()gates both entry points. It is deliberately stricter than the existingvalid(): the offset is the low 16 bits of an id and so can name a slot pastMAX_SIZEin a blob that is full, and the next free slot is not allocated yet. That second case is not theoretical — a test that marked the free slot caused an unrelated metric created later in the same run to come out invisible, becausecreate()only clears the flag when it finds the name already present, not when it allocates a fresh slot.A
static_assertnow tiesMAX_BLOBSandMAX_SIZEto the widths of the blob and offset fields in an id. That relationship is what keeps every_blobs[]subscript in this class in range without an explicit check, and nothing previously enforced it.API change
Metrics::iterator,begin(),end()andfind()are removed.Metrics.his inTSUTIL_PUBLIC_HEADERSand is installed, so this is a source level break for anything enumerating a store directly or constructing an iterator.Nothing in tree does either, and
find()had no callers even in tests before this PR added some. Worth notingfind()predated this PR and was safe there — the positional constructor did not skip, so equality was a plain position compare and no bound was involved. Unlisting is what made it unsound, which is the argument for removing it here rather than in a separate change.Worth a line in the 11.0.0 release notes alongside the
createSpanandrenameremovals from #13583.Tests
test_Metrics.cc, inTEST_CASE("Metrics unlisting"): skipped by enumeration; relisted bycreate()with its value intact; unlist and relist by name; still resolvable by name and id while unlisted;for_eachskipping an unlisted first slot; an unlisted run at the end of the store; the reported type matching the type each metric was created with; three rejected id shapes (an unallocated blob, an offset pastMAX_SIZE, and the next free slot); and independence between the published and hidden stores.TEST_CASE("Metrics")coversfor_eachitself: the reserved slot first, creation order, and the count moving by one when a metric is created.test_RecHiddenMetricLookup.cc: an unlisted metric is not enumerated byRecLookupMatchingRecords, is still found byRecLookupRecord, and returns to enumeration when relisted.Documented in
doc/developer-guide/internal-libraries/Metrics.en.rst, which now has anEnumerating metricssection stating thatfor_eachis the only way to enumerate and why.