Skip to content

Thread filter optim - #238

Merged
r1viollet merged 20 commits into
mainfrom
r1viollet/thread_filter_squash
Sep 11, 2025
Merged

Thread filter optim#238
r1viollet merged 20 commits into
mainfrom
r1viollet/thread_filter_squash

Conversation

@r1viollet

@r1violletr1viollet commented Jul 7, 2025

Copy link
Copy Markdown
Contributor

What does this PR do?:

  • Reserve padded slots
  • Introduce a register / unregister to retrieve slots
  • manage a free list

Motivation:

Improve throughput of applications that run on many threads with many context updates.

Additional Notes:

How to test the change?:

For Datadog employees:

  • If this PR touches code that signs or publishes builds or packages, or handles
    credentials of any kind, I've requested a review from @DataDog/security-design-and-guidance.
  • This PR doesn't touch any of that.
  • JIRA: [JIRA-XXXX]

Unsure? Have a question? Request a review!

@github-actions

github-actionsBot commented Jul 7, 2025

Copy link
Copy Markdown
Contributor

🔧 Report generated by pr-comment-cppcheck

CppCheck Report

Errors (2)

Warnings (8)

Style Violations (306)

@github-actions

github-actionsBot commented Jul 7, 2025

Copy link
Copy Markdown
Contributor

🔧 Report generated by pr-comment-scanbuild

@r1viollet
r1violletforce-pushed the r1viollet/thread_filter_squash branch 3 times, most recently from e5bce28 to 0918008CompareJuly 7, 2025 11:45
@r1viollet

Copy link
Copy Markdown
ContributorAuthor

I have reasonable performance on most runs:

Benchmark (command) (skipResults) (workload) Mode Cnt Score Error Units
ThreadFilterBenchmark.threadFilterStress01 cpu=100us,wall=100us,filter=1 true 0 avgt 0.039 us/op
ThreadFilterBenchmark.threadFilterStress01 cpu=100us,wall=100us,filter=1 true 7 avgt 0.041 us/op
ThreadFilterBenchmark.threadFilterStress01 cpu=100us,wall=100us,filter=1 true 70000 avgt 111.094 us/op
ThreadFilterBenchmark.threadFilterStress02 cpu=100us,wall=100us,filter=1 true 0 avgt 0.132 us/op
ThreadFilterBenchmark.threadFilterStress02 cpu=100us,wall=100us,filter=1 true 7 avgt 0.139 us/op
ThreadFilterBenchmark.threadFilterStress02 cpu=100us,wall=100us,filter=1 true 70000 avgt 108.666 us/op
ThreadFilterBenchmark.threadFilterStress04 cpu=100us,wall=100us,filter=1 true 0 avgt 0.258 us/op
ThreadFilterBenchmark.threadFilterStress04 cpu=100us,wall=100us,filter=1 true 7 avgt 0.278 us/op
ThreadFilterBenchmark.threadFilterStress04 cpu=100us,wall=100us,filter=1 true 70000 avgt 118.940 us/op
ThreadFilterBenchmark.threadFilterStress08 cpu=100us,wall=100us,filter=1 true 0 avgt 0.624 us/op
ThreadFilterBenchmark.threadFilterStress08 cpu=100us,wall=100us,filter=1 true 7 avgt 0.646 us/op
ThreadFilterBenchmark.threadFilterStress08 cpu=100us,wall=100us,filter=1 true 70000 avgt 160.170 us/op
ThreadFilterBenchmark.threadFilterStress16 cpu=100us,wall=100us,filter=1 true 0 avgt 1.780 us/op
ThreadFilterBenchmark.threadFilterStress16 cpu=100us,wall=100us,filter=1 true 7 avgt 2.288 us/op
ThreadFilterBenchmark.threadFilterStress16 cpu=100us,wall=100us,filter=1 true 70000 avgt 221.987 us/op

I'm not sure why some runs still blow up for higher numbers of threads.

@r1violletr1viollet mentioned this pull request Jul 7, 2025
3 tasks
@jbachorik
jbachorikforce-pushed the r1viollet/thread_filter_squash branch 2 times, most recently from e0ac246 to 2421ba9CompareJuly 10, 2025 12:48
Comment threadddprof-lib/src/main/cpp/profiler.cpp
Comment threadddprof-lib/src/main/cpp/threadFilter.cpp
@r1viollet

r1viollet commented Jul 10, 2025

Copy link
Copy Markdown
ContributorAuthor

CppCheck Report

Errors (2)

Warnings (8)

Style Violations (305)

@jbachorik
jbachorikforce-pushed the r1viollet/thread_filter_squash branch from 2421ba9 to 50a8d5fCompareJuly 21, 2025 14:57
@jbachorik

Copy link
Copy Markdown
Collaborator

I did run some comparison of native memory usage with different thread filter implementations - data is in the notebook

TL;DR there is no observable increase in the native memory usage (the UNDEFINED category). Anyway, it would be useful to have an extra counter for the ThreadIDTable utilization.

jbachorikand others added 3 commits July 24, 2025 21:43
If the TLS cleanup fires before the JVMTI hook, we want to
ensure that we don't crash while retrieving the ProfiledThread
- Add a check on validity of ProfiledThread
Comment threadddprof-lib/src/main/cpp/threadFilter.cpp
Comment threadddprof-lib/src/main/cpp/threadFilter.cpp
Comment threadddprof-lib/src/main/cpp/threadFilter.cpp Outdated
- Start the profiler to ensure we have valid thread objects
- add asserts around missing thread object
- remove print (replacing with an assert)
@r1viollet

r1viollet commented Aug 21, 2025

Copy link
Copy Markdown
ContributorAuthor

CppCheck Report

Errors (2)

Warnings (8)

Style Violations (305)

@r1viollet
r1viollet marked this pull request as ready for review August 21, 2025 07:53
@r1viollet
r1violletforce-pushed the r1viollet/thread_filter_squash branch from 6171739 to e30d88fCompareAugust 21, 2025 14:05
- Fix removal of self in timerloop init
it was not using a slotID but a thread ID
- Add assertion to find other potential issues
@r1viollet
r1violletforce-pushed the r1viollet/thread_filter_squash branch from e30d88f to e78a6b2CompareAugust 21, 2025 14:08
Comment threadddprof-lib/src/main/cpp/javaApi.cpp
Java_com_datadoghq_profiler_JavaProfiler_filterThreadRemove0(JNIEnv *env,
jobject unused) {
ProfiledThread *current = ProfiledThread::current();
if (unlikely(current == nullptr)) {

@zhengyu123zhengyu123Aug 21, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think assert(current != nllptr) should be sufficient, otherwise, we have a bigger problem.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this happens on unloading. JVMTI cleanup can be removed before all threads are finished ?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have the assert for debug builds, though we can keep avoiding crashes for release builds. Feel free to answer if you do not agree.

return;
}
int tid = current->tid();
if (unlikely(tid < 0)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible? or we should just assert

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question
I think we are missing the instrumentation of some threads. Could non-java threads call back into java?
It could be nice to have asserts to debug this, though I think I'd prefer the safer path for a prod release.

Comment threadddprof-lib/src/main/cpp/javaApi.cpp
return;
}
int tid = current->tid();
if (unlikely(tid < 0)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as above

private native void stop0() throws IllegalStateException;
private native String execute0(String command) throws IllegalArgumentException, IllegalStateException, IOException;

private native void filterThreadAdd0();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like that filterThreadAdd0 == filterThread0(true) and filterThreadRemove0 == filterThread0(false). Please remove duplications.

void collect(std::vector<int> &v);
private:
// Optimized slot structure with padding to avoid false sharing
struct alignas(64) Slot {

@zhengyu123zhengyu123Aug 21, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have definition of DEFAULT_CACHE_LINE_SIZE in dd_arch.h. I would suggest following code for readability and portability.

 struct alignas(DEFAULT_CACHE_LINE_SIZE) Slot {
std::atomic<int> value{-1};
char padding[DEFAULT_CACHE_LINE_SIZE - sizeof(value)];
};

std::atomic<SlotID> _next_index{0};
std::unique_ptr<FreeListNode[]> _free_list;

struct alignas(64) ShardHead { std::atomic<int> head{-1}; };

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use DEFAULT_CACHE_LINE_SIZE for readability and portability.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks

}
// Try to install it atomically
ChunkStorage* expected = nullptr;
if (_chunks[chunk_idx].compare_exchange_strong(expected, new_chunk, std::memory_order_acq_rel)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

memory_order_release should be sufficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

agreed


ThreadFilter::SlotID ThreadFilter::registerThread() {
// If disabled, block new registrations
if (!_enabled.load(std::memory_order_acquire)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't see any memory ordering _enabled providing. Could you explain what it releases

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We want the init to be called (so that constructor is finished)
This means other threads might not have the init finished
Though we can always lazily register them later (even if the on thread start path is better)
So this acts as a load barrier. I think it makes sense. Feel free to challenge.

return;
}

ChunkStorage* chunk = _chunks[chunk_idx].load(std::memory_order_relaxed);

@zhengyu123zhengyu123Aug 21, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you need memory_order_acquire ordering here to match the release store.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same as the add, I think we can keep relaxed. Feel free to challenge.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you do need std::memory_order_acquire, just as you did in ThreadFilter::accept()

@zhengyu123zhengyu123 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did partial second round reviewing, I think there are many inconsistencies in memory ordering.

_num_chunks.store(0, std::memory_order_release);
// Detach and delete chunks
for (int i = 0; i < kMaxChunks; ++i) {
ChunkStorage* chunk = _chunks[i].exchange(nullptr, std::memory_order_acq_rel);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

memory_order_acquire instead of memory_order_acq_rel

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

adjusting

int slot_idx = slot_id & kChunkMask;

// Fast path: assume valid slot_id from registerThread()
ChunkStorage* chunk = _chunks[chunk_idx].load(std::memory_order_relaxed);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Need memory_order_acquire ordering

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was intentional, the idea is that we either have null, or a valid chunk. The impact of missing one trace is acceptable if we don't see the chunk in the current thread (which is pretty unlikely). wdyt ?

// Fast path: assume valid slot_id from registerThread()
ChunkStorage* chunk = _chunks[chunk_idx].load(std::memory_order_relaxed);
if (likely(chunk != nullptr)) {
return chunk->slots[slot_idx].value.load(std::memory_order_acquire) != -1;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

memory_order_relaxed should be sufficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is a tradeoff. I hope this is a path where acquire is acceptable. If you think otherwise I can adjust.

@r1viollet

Copy link
Copy Markdown
ContributorAuthor

Thanks @zhengyu123 I really appreciate the thorough review. Apologies if I only have a little time to spend on this every week.

…t/thread_filter_squash
In the current merge, I'm removing the active bitmap. The ActiveBitmap needs to be adjusted to the slot logics.
We can adjust this by retrieving the address of the slot.
This can be simpler than with the bitmap.
Comment threadddprof-lib/src/main/cpp/javaApi.cpp
Comment threadddprof-lib/src/main/cpp/javaApi.cpp
int slot_idx = slot_id & kChunkMask;

// Fast path: assume valid slot_id from registerThread()
ChunkStorage* chunk = _chunks[chunk_idx].load(std::memory_order_relaxed);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you do need std::memory_order_acquire, just as you did in ThreadFilter::accept() and match the release from ChunkStorage* expected = nullptr; if (_chunks[chunk_idx].compare_exchange_strong(expected, new_chunk, std::memory_order_release))

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure we do, but let's measure before we debate. I'll run the test to see if this changes anything on perf.

return;
}

ChunkStorage* chunk = _chunks[chunk_idx].load(std::memory_order_relaxed);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you do need std::memory_order_acquire, just as you did in ThreadFilter::accept()

@r1viollet

Copy link
Copy Markdown
ContributorAuthor

@zhengyu123 are we OK with this version ?

// Collect thread IDs from the fixed-size table into the main set
_thread_ids[i][old_index].collect(threads);
_thread_ids[i][old_index].clear();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Release here?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think memory_order_acq_rel is good here for the old_index. Though we might be talking about something else.

ProfiledThread *current = ProfiledThread::current();
int tid = -1;

if (current != nullptr) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can current == nullptr?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I remember seeing a crash around this..
Basically JVMTI could unload and delete the profiled thread from under our feet.

@r1viollet
r1viollet merged commit 41fcf55 into mainSep 11, 2025
94 checks passed
@r1viollet
r1viollet deleted the r1viollet/thread_filter_squash branch September 11, 2025 14:49
@github-actionsgithub-actionsBot added this to the 1.32.0 milestone Sep 11, 2025
@r1viollet

Copy link
Copy Markdown
ContributorAuthor

Notes for future: if we care about the ~100KB, we can reduce the padding. This becomes a tradeoff between memory and risk of false sharing.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@r1viollet@jbachorik@zhengyu123