Fix deadlock when concurrent tasks queue the same asset's downstream dags - #70128

Closed
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock
Closed

Fix deadlock when concurrent tasks queue the same asset's downstream dags#70128
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock

Conversation

@Andrushika

Copy link
Copy Markdown
Contributor

Fix deadlock when concurrent tasks queue the same asset's downstream dags

Background

When a task produces an asset, the downstream Dags scheduled on that asset need to run. Airflow records this in the asset_dag_run_queue (ADRQ) table, one row per (asset_id, target_dag_id). So when a task succeeds and updates asset S, register_asset_change inserts one ADRQ row for every Dag scheduled on S. Those Dags are collected into a Python set (dags_to_queue), and the rows are inserted by looping over that set.

Why

A Python set has no stable order (The essence of set is a hash table). Here, the elements are DagModel objects, freshly loaded at different memory addresses in each process, so two processes can loop over the same Dags in opposite orders.

When two tasks finish at the same time and both emit to the same asset, they insert the same ADRQ rows, but maybe in opposite orders. Each insert holds a row lock until commit. So transaction A can hold row (S, dag_a) and wait for (S, dag_b), while transaction B holds (S, dag_b) and waits for (S, dag_a). That is a deadlock. Postgres aborts one side and it retries, which adds latency and error noise under high asset fan-out.

What

Insert the rows in a fixed order, sorted by dag_id, so every process takes the row locks in the same order and the cycle cannot form. A small shared helper _sorted_by_dag_id is used in all three insert paths (postgres ON CONFLICT, mysql ON DUPLICATE KEY, and the per-row SAVEPOINT fallback), because all three loop over the same set.

I was running the concurrent test for #70078 and found that this problem exists on main.
Then I reproduced it on Postgres with a script that has 8 concurrent transactions insert one asset's downstream rows in opposite orders (what two processes' set iteration can produce). Over 20 rounds it hit 658 deadlocks with the old order and 0 after sorting. I only benchmarked Postgres. MySQL uses InnoDB row locks with the same inversion, so sorting fixes it there too. The SQLite path cannot deadlock this way and is sorted only for consistency.

The script for reproduce:
from __future__ importannotationsimportthreadingimporttimefromsqlalchemyimportdeletefromsqlalchemy.dialects.postgresqlimportinsertfromsqlalchemy.excimportOperationalErrorfromairflowimportsettingsfromairflow.models.assetimportAssetActive, AssetDagRunQueue, AssetModel, DagScheduleAssetReferencefromairflow.models.dagimportDagModelfromairflow.models.dagbundleimportDagBundleModelfromairflow.utils.sessionimportcreate_sessionASSET_ID=1DOWNSTREAM_DAGS= ["dag_a", "dag_b", "dag_c", "dag_d"]
N_THREADS=8N_ROUNDS=20ROW_GAP=0.004# small pause between rows to widen the lock windowdef_seed():
withcreate_session() ass:
s.execute(delete(AssetDagRunQueue))
s.execute(delete(DagScheduleAssetReference))
s.query(AssetActive).delete()
s.query(AssetModel).delete()
s.query(DagModel).filter(DagModel.dag_id.in_(DOWNSTREAM_DAGS)).delete(synchronize_session=False)
s.merge(DagBundleModel(name="repro"))
s.flush() # bundle must exist before dags reference it (FK)asset=AssetModel(id=ASSET_ID, name="S", uri="s3://bucket/S", group="asset", extra={})
s.add_all([asset, AssetActive.for_asset(asset)])
fordinDOWNSTREAM_DAGS:
s.add(DagModel(dag_id=d, bundle_name="repro", is_stale=False, fileloc=f"{d}.py"))
s.flush()
asset.scheduled_dags= [DagScheduleAssetReference(dag_id=d) fordinDOWNSTREAM_DAGS]
def_insert_rows(order, sort_fix, counters, barrier):
dag_ids=sorted(order) ifsort_fixelseorder# the fix: always sort firstbarrier.wait()
whileTrue:
try:
withcreate_session() assession:
fordag_idindag_ids:
stmt= (
insert(AssetDagRunQueue)
.values(asset_id=ASSET_ID, target_dag_id=dag_id)
.on_conflict_do_nothing()
)
session.execute(stmt)
time.sleep(ROW_GAP)
returnexceptOperationalErrorase:
if"deadlock detected"instr(e).lower():
counters["deadlocks"] +=1continue# retry: the losing side re-runs and eventually winsraisedef_run(sort_fix):
counters= {"deadlocks": 0}
forward=DOWNSTREAM_DAGSreverse=list(reversed(DOWNSTREAM_DAGS))
for_inrange(N_ROUNDS):
withcreate_session() ass: # fresh rows each round so ON CONFLICT actually insertss.execute(delete(AssetDagRunQueue))
barrier=threading.Barrier(N_THREADS)
threads= [
threading.Thread(
target=_insert_rows,
args=(forwardifi%2==0elsereverse, sort_fix, counters, barrier),
)
foriinrange(N_THREADS)
]
fortinthreads:
t.start()
fortinthreads:
t.join()
returncounters["deadlocks"]
defmain():
ifsettings.engine.dialect.name!="postgresql":
raiseSystemExit("Run under --backend postgres (row locks are needed to reproduce).")
_seed()
unsorted_deadlocks=_run(sort_fix=False)
sorted_deadlocks=_run(sort_fix=True)
print(
f"\n=== ADRQ insert-order deadlock — {N_THREADS} threads x {N_ROUNDS} rounds, "f"{len(DOWNSTREAM_DAGS)} downstream dags ===\n"f" UNSORTED (insert in set order, forward vs reversed): {unsorted_deadlocks} deadlocks\n"f" SORTED (sort by dag_id first, the fix) : {sorted_deadlocks} deadlocks\n"
)
if__name__=="__main__":
main()
  • Yes (please specify the tool below)

Generated-by: Claude Code Opus 4.8 following the guidelines


  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

…dags
When a task updates an asset, register_asset_change queues an
AssetDagRunQueue row for each of the asset's downstream dags. The dags
come from a set, whose iteration order varies between processes, and the
rows were inserted in that order. Two tasks completing at once and
emitting to the same asset could take the per-row locks in opposite
orders and deadlock. Insert in a fixed dag_id order so the lock order is
consistent across processes; this covers the postgres, mysql, and
per-row SAVEPOINT paths.
@Andrushika
Andrushika marked this pull request as ready for review July 20, 2026 13:06
Comment threadairflow-core/src/airflow/assets/manager.py
Comment threadairflow-core/src/airflow/assets/manager.py Outdated
@Andrushika

Copy link
Copy Markdown
ContributorAuthor

After #70972, this deadlock issue does not exist anymore. Closing it.
@Dev-iL, big thanks for your review!

@Andrushika
Andrushika deleted the fix-asset-dag-run-queue-insert-deadlock branch August 7, 2026 05:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Andrushika@Dev-iL
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Fix deadlock when concurrent tasks queue the same asset's downstream dags - #70128

Closed
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock
Closed

Fix deadlock when concurrent tasks queue the same asset's downstream dags#70128
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock

Conversation

@Andrushika

Copy link
Copy Markdown
Contributor

Fix deadlock when concurrent tasks queue the same asset's downstream dags

Background

When a task produces an asset, the downstream Dags scheduled on that asset need to run. Airflow records this in the asset_dag_run_queue (ADRQ) table, one row per (asset_id, target_dag_id). So when a task succeeds and updates asset S, register_asset_change inserts one ADRQ row for every Dag scheduled on S. Those Dags are collected into a Python set (dags_to_queue), and the rows are inserted by looping over that set.

Why

A Python set has no stable order (The essence of set is a hash table). Here, the elements are DagModel objects, freshly loaded at different memory addresses in each process, so two processes can loop over the same Dags in opposite orders.

When two tasks finish at the same time and both emit to the same asset, they insert the same ADRQ rows, but maybe in opposite orders. Each insert holds a row lock until commit. So transaction A can hold row (S, dag_a) and wait for (S, dag_b), while transaction B holds (S, dag_b) and waits for (S, dag_a). That is a deadlock. Postgres aborts one side and it retries, which adds latency and error noise under high asset fan-out.

What

Insert the rows in a fixed order, sorted by dag_id, so every process takes the row locks in the same order and the cycle cannot form. A small shared helper _sorted_by_dag_id is used in all three insert paths (postgres ON CONFLICT, mysql ON DUPLICATE KEY, and the per-row SAVEPOINT fallback), because all three loop over the same set.

I was running the concurrent test for #70078 and found that this problem exists on main.
Then I reproduced it on Postgres with a script that has 8 concurrent transactions insert one asset's downstream rows in opposite orders (what two processes' set iteration can produce). Over 20 rounds it hit 658 deadlocks with the old order and 0 after sorting. I only benchmarked Postgres. MySQL uses InnoDB row locks with the same inversion, so sorting fixes it there too. The SQLite path cannot deadlock this way and is sorted only for consistency.

The script for reproduce:
from __future__ importannotationsimportthreadingimporttimefromsqlalchemyimportdeletefromsqlalchemy.dialects.postgresqlimportinsertfromsqlalchemy.excimportOperationalErrorfromairflowimportsettingsfromairflow.models.assetimportAssetActive, AssetDagRunQueue, AssetModel, DagScheduleAssetReferencefromairflow.models.dagimportDagModelfromairflow.models.dagbundleimportDagBundleModelfromairflow.utils.sessionimportcreate_sessionASSET_ID=1DOWNSTREAM_DAGS= ["dag_a", "dag_b", "dag_c", "dag_d"]
N_THREADS=8N_ROUNDS=20ROW_GAP=0.004# small pause between rows to widen the lock windowdef_seed():
withcreate_session() ass:
s.execute(delete(AssetDagRunQueue))
s.execute(delete(DagScheduleAssetReference))
s.query(AssetActive).delete()
s.query(AssetModel).delete()
s.query(DagModel).filter(DagModel.dag_id.in_(DOWNSTREAM_DAGS)).delete(synchronize_session=False)
s.merge(DagBundleModel(name="repro"))
s.flush() # bundle must exist before dags reference it (FK)asset=AssetModel(id=ASSET_ID, name="S", uri="s3://bucket/S", group="asset", extra={})
s.add_all([asset, AssetActive.for_asset(asset)])
fordinDOWNSTREAM_DAGS:
s.add(DagModel(dag_id=d, bundle_name="repro", is_stale=False, fileloc=f"{d}.py"))
s.flush()
asset.scheduled_dags= [DagScheduleAssetReference(dag_id=d) fordinDOWNSTREAM_DAGS]
def_insert_rows(order, sort_fix, counters, barrier):
dag_ids=sorted(order) ifsort_fixelseorder# the fix: always sort firstbarrier.wait()
whileTrue:
try:
withcreate_session() assession:
fordag_idindag_ids:
stmt= (
insert(AssetDagRunQueue)
.values(asset_id=ASSET_ID, target_dag_id=dag_id)
.on_conflict_do_nothing()
)
session.execute(stmt)
time.sleep(ROW_GAP)
returnexceptOperationalErrorase:
if"deadlock detected"instr(e).lower():
counters["deadlocks"] +=1continue# retry: the losing side re-runs and eventually winsraisedef_run(sort_fix):
counters= {"deadlocks": 0}
forward=DOWNSTREAM_DAGSreverse=list(reversed(DOWNSTREAM_DAGS))
for_inrange(N_ROUNDS):
withcreate_session() ass: # fresh rows each round so ON CONFLICT actually insertss.execute(delete(AssetDagRunQueue))
barrier=threading.Barrier(N_THREADS)
threads= [
threading.Thread(
target=_insert_rows,
args=(forwardifi%2==0elsereverse, sort_fix, counters, barrier),
)
foriinrange(N_THREADS)
]
fortinthreads:
t.start()
fortinthreads:
t.join()
returncounters["deadlocks"]
defmain():
ifsettings.engine.dialect.name!="postgresql":
raiseSystemExit("Run under --backend postgres (row locks are needed to reproduce).")
_seed()
unsorted_deadlocks=_run(sort_fix=False)
sorted_deadlocks=_run(sort_fix=True)
print(
f"\n=== ADRQ insert-order deadlock — {N_THREADS} threads x {N_ROUNDS} rounds, "f"{len(DOWNSTREAM_DAGS)} downstream dags ===\n"f" UNSORTED (insert in set order, forward vs reversed): {unsorted_deadlocks} deadlocks\n"f" SORTED (sort by dag_id first, the fix) : {sorted_deadlocks} deadlocks\n"
)
if__name__=="__main__":
main()
  • Yes (please specify the tool below)

Generated-by: Claude Code Opus 4.8 following the guidelines


  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

…dags
When a task updates an asset, register_asset_change queues an
AssetDagRunQueue row for each of the asset's downstream dags. The dags
come from a set, whose iteration order varies between processes, and the
rows were inserted in that order. Two tasks completing at once and
emitting to the same asset could take the per-row locks in opposite
orders and deadlock. Insert in a fixed dag_id order so the lock order is
consistent across processes; this covers the postgres, mysql, and
per-row SAVEPOINT paths.
@Andrushika
Andrushika marked this pull request as ready for review July 20, 2026 13:06
Comment threadairflow-core/src/airflow/assets/manager.py
Comment threadairflow-core/src/airflow/assets/manager.py Outdated
@Andrushika

Copy link
Copy Markdown
ContributorAuthor

After #70972, this deadlock issue does not exist anymore. Closing it.
@Dev-iL, big thanks for your review!

@Andrushika
Andrushika deleted the fix-asset-dag-run-queue-insert-deadlock branch August 7, 2026 05:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Andrushika@Dev-iL
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix deadlock when concurrent tasks queue the same asset's downstream dags - #70128

Closed
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock
Closed

Fix deadlock when concurrent tasks queue the same asset's downstream dags#70128
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock

Conversation

@Andrushika

Copy link
Copy Markdown
Contributor

Fix deadlock when concurrent tasks queue the same asset's downstream dags

Background

When a task produces an asset, the downstream Dags scheduled on that asset need to run. Airflow records this in the asset_dag_run_queue (ADRQ) table, one row per (asset_id, target_dag_id). So when a task succeeds and updates asset S, register_asset_change inserts one ADRQ row for every Dag scheduled on S. Those Dags are collected into a Python set (dags_to_queue), and the rows are inserted by looping over that set.

Why

A Python set has no stable order (The essence of set is a hash table). Here, the elements are DagModel objects, freshly loaded at different memory addresses in each process, so two processes can loop over the same Dags in opposite orders.

When two tasks finish at the same time and both emit to the same asset, they insert the same ADRQ rows, but maybe in opposite orders. Each insert holds a row lock until commit. So transaction A can hold row (S, dag_a) and wait for (S, dag_b), while transaction B holds (S, dag_b) and waits for (S, dag_a). That is a deadlock. Postgres aborts one side and it retries, which adds latency and error noise under high asset fan-out.

What

Insert the rows in a fixed order, sorted by dag_id, so every process takes the row locks in the same order and the cycle cannot form. A small shared helper _sorted_by_dag_id is used in all three insert paths (postgres ON CONFLICT, mysql ON DUPLICATE KEY, and the per-row SAVEPOINT fallback), because all three loop over the same set.

I was running the concurrent test for #70078 and found that this problem exists on main.
Then I reproduced it on Postgres with a script that has 8 concurrent transactions insert one asset's downstream rows in opposite orders (what two processes' set iteration can produce). Over 20 rounds it hit 658 deadlocks with the old order and 0 after sorting. I only benchmarked Postgres. MySQL uses InnoDB row locks with the same inversion, so sorting fixes it there too. The SQLite path cannot deadlock this way and is sorted only for consistency.

The script for reproduce:
from __future__ importannotationsimportthreadingimporttimefromsqlalchemyimportdeletefromsqlalchemy.dialects.postgresqlimportinsertfromsqlalchemy.excimportOperationalErrorfromairflowimportsettingsfromairflow.models.assetimportAssetActive, AssetDagRunQueue, AssetModel, DagScheduleAssetReferencefromairflow.models.dagimportDagModelfromairflow.models.dagbundleimportDagBundleModelfromairflow.utils.sessionimportcreate_sessionASSET_ID=1DOWNSTREAM_DAGS= ["dag_a", "dag_b", "dag_c", "dag_d"]
N_THREADS=8N_ROUNDS=20ROW_GAP=0.004# small pause between rows to widen the lock windowdef_seed():
withcreate_session() ass:
s.execute(delete(AssetDagRunQueue))
s.execute(delete(DagScheduleAssetReference))
s.query(AssetActive).delete()
s.query(AssetModel).delete()
s.query(DagModel).filter(DagModel.dag_id.in_(DOWNSTREAM_DAGS)).delete(synchronize_session=False)
s.merge(DagBundleModel(name="repro"))
s.flush() # bundle must exist before dags reference it (FK)asset=AssetModel(id=ASSET_ID, name="S", uri="s3://bucket/S", group="asset", extra={})
s.add_all([asset, AssetActive.for_asset(asset)])
fordinDOWNSTREAM_DAGS:
s.add(DagModel(dag_id=d, bundle_name="repro", is_stale=False, fileloc=f"{d}.py"))
s.flush()
asset.scheduled_dags= [DagScheduleAssetReference(dag_id=d) fordinDOWNSTREAM_DAGS]
def_insert_rows(order, sort_fix, counters, barrier):
dag_ids=sorted(order) ifsort_fixelseorder# the fix: always sort firstbarrier.wait()
whileTrue:
try:
withcreate_session() assession:
fordag_idindag_ids:
stmt= (
insert(AssetDagRunQueue)
.values(asset_id=ASSET_ID, target_dag_id=dag_id)
.on_conflict_do_nothing()
)
session.execute(stmt)
time.sleep(ROW_GAP)
returnexceptOperationalErrorase:
if"deadlock detected"instr(e).lower():
counters["deadlocks"] +=1continue# retry: the losing side re-runs and eventually winsraisedef_run(sort_fix):
counters= {"deadlocks": 0}
forward=DOWNSTREAM_DAGSreverse=list(reversed(DOWNSTREAM_DAGS))
for_inrange(N_ROUNDS):
withcreate_session() ass: # fresh rows each round so ON CONFLICT actually insertss.execute(delete(AssetDagRunQueue))
barrier=threading.Barrier(N_THREADS)
threads= [
threading.Thread(
target=_insert_rows,
args=(forwardifi%2==0elsereverse, sort_fix, counters, barrier),
)
foriinrange(N_THREADS)
]
fortinthreads:
t.start()
fortinthreads:
t.join()
returncounters["deadlocks"]
defmain():
ifsettings.engine.dialect.name!="postgresql":
raiseSystemExit("Run under --backend postgres (row locks are needed to reproduce).")
_seed()
unsorted_deadlocks=_run(sort_fix=False)
sorted_deadlocks=_run(sort_fix=True)
print(
f"\n=== ADRQ insert-order deadlock — {N_THREADS} threads x {N_ROUNDS} rounds, "f"{len(DOWNSTREAM_DAGS)} downstream dags ===\n"f" UNSORTED (insert in set order, forward vs reversed): {unsorted_deadlocks} deadlocks\n"f" SORTED (sort by dag_id first, the fix) : {sorted_deadlocks} deadlocks\n"
)
if__name__=="__main__":
main()
  • Yes (please specify the tool below)

Generated-by: Claude Code Opus 4.8 following the guidelines


  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

…dags
When a task updates an asset, register_asset_change queues an
AssetDagRunQueue row for each of the asset's downstream dags. The dags
come from a set, whose iteration order varies between processes, and the
rows were inserted in that order. Two tasks completing at once and
emitting to the same asset could take the per-row locks in opposite
orders and deadlock. Insert in a fixed dag_id order so the lock order is
consistent across processes; this covers the postgres, mysql, and
per-row SAVEPOINT paths.
@Andrushika
Andrushika marked this pull request as ready for review July 20, 2026 13:06
Comment threadairflow-core/src/airflow/assets/manager.py
Comment threadairflow-core/src/airflow/assets/manager.py Outdated
@Andrushika

Copy link
Copy Markdown
ContributorAuthor

After #70972, this deadlock issue does not exist anymore. Closing it.
@Dev-iL, big thanks for your review!

@Andrushika
Andrushika deleted the fix-asset-dag-run-queue-insert-deadlock branch August 7, 2026 05:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Andrushika@Dev-iL
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix deadlock when concurrent tasks queue the same asset's downstream dags - #70128

Closed
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock
Closed

Fix deadlock when concurrent tasks queue the same asset's downstream dags#70128
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock

Conversation

@Andrushika

Copy link
Copy Markdown
Contributor

Fix deadlock when concurrent tasks queue the same asset's downstream dags

Background

When a task produces an asset, the downstream Dags scheduled on that asset need to run. Airflow records this in the asset_dag_run_queue (ADRQ) table, one row per (asset_id, target_dag_id). So when a task succeeds and updates asset S, register_asset_change inserts one ADRQ row for every Dag scheduled on S. Those Dags are collected into a Python set (dags_to_queue), and the rows are inserted by looping over that set.

Why

A Python set has no stable order (The essence of set is a hash table). Here, the elements are DagModel objects, freshly loaded at different memory addresses in each process, so two processes can loop over the same Dags in opposite orders.

When two tasks finish at the same time and both emit to the same asset, they insert the same ADRQ rows, but maybe in opposite orders. Each insert holds a row lock until commit. So transaction A can hold row (S, dag_a) and wait for (S, dag_b), while transaction B holds (S, dag_b) and waits for (S, dag_a). That is a deadlock. Postgres aborts one side and it retries, which adds latency and error noise under high asset fan-out.

What

Insert the rows in a fixed order, sorted by dag_id, so every process takes the row locks in the same order and the cycle cannot form. A small shared helper _sorted_by_dag_id is used in all three insert paths (postgres ON CONFLICT, mysql ON DUPLICATE KEY, and the per-row SAVEPOINT fallback), because all three loop over the same set.

I was running the concurrent test for #70078 and found that this problem exists on main.
Then I reproduced it on Postgres with a script that has 8 concurrent transactions insert one asset's downstream rows in opposite orders (what two processes' set iteration can produce). Over 20 rounds it hit 658 deadlocks with the old order and 0 after sorting. I only benchmarked Postgres. MySQL uses InnoDB row locks with the same inversion, so sorting fixes it there too. The SQLite path cannot deadlock this way and is sorted only for consistency.

The script for reproduce:
from __future__ importannotationsimportthreadingimporttimefromsqlalchemyimportdeletefromsqlalchemy.dialects.postgresqlimportinsertfromsqlalchemy.excimportOperationalErrorfromairflowimportsettingsfromairflow.models.assetimportAssetActive, AssetDagRunQueue, AssetModel, DagScheduleAssetReferencefromairflow.models.dagimportDagModelfromairflow.models.dagbundleimportDagBundleModelfromairflow.utils.sessionimportcreate_sessionASSET_ID=1DOWNSTREAM_DAGS= ["dag_a", "dag_b", "dag_c", "dag_d"]
N_THREADS=8N_ROUNDS=20ROW_GAP=0.004# small pause between rows to widen the lock windowdef_seed():
withcreate_session() ass:
s.execute(delete(AssetDagRunQueue))
s.execute(delete(DagScheduleAssetReference))
s.query(AssetActive).delete()
s.query(AssetModel).delete()
s.query(DagModel).filter(DagModel.dag_id.in_(DOWNSTREAM_DAGS)).delete(synchronize_session=False)
s.merge(DagBundleModel(name="repro"))
s.flush() # bundle must exist before dags reference it (FK)asset=AssetModel(id=ASSET_ID, name="S", uri="s3://bucket/S", group="asset", extra={})
s.add_all([asset, AssetActive.for_asset(asset)])
fordinDOWNSTREAM_DAGS:
s.add(DagModel(dag_id=d, bundle_name="repro", is_stale=False, fileloc=f"{d}.py"))
s.flush()
asset.scheduled_dags= [DagScheduleAssetReference(dag_id=d) fordinDOWNSTREAM_DAGS]
def_insert_rows(order, sort_fix, counters, barrier):
dag_ids=sorted(order) ifsort_fixelseorder# the fix: always sort firstbarrier.wait()
whileTrue:
try:
withcreate_session() assession:
fordag_idindag_ids:
stmt= (
insert(AssetDagRunQueue)
.values(asset_id=ASSET_ID, target_dag_id=dag_id)
.on_conflict_do_nothing()
)
session.execute(stmt)
time.sleep(ROW_GAP)
returnexceptOperationalErrorase:
if"deadlock detected"instr(e).lower():
counters["deadlocks"] +=1continue# retry: the losing side re-runs and eventually winsraisedef_run(sort_fix):
counters= {"deadlocks": 0}
forward=DOWNSTREAM_DAGSreverse=list(reversed(DOWNSTREAM_DAGS))
for_inrange(N_ROUNDS):
withcreate_session() ass: # fresh rows each round so ON CONFLICT actually insertss.execute(delete(AssetDagRunQueue))
barrier=threading.Barrier(N_THREADS)
threads= [
threading.Thread(
target=_insert_rows,
args=(forwardifi%2==0elsereverse, sort_fix, counters, barrier),
)
foriinrange(N_THREADS)
]
fortinthreads:
t.start()
fortinthreads:
t.join()
returncounters["deadlocks"]
defmain():
ifsettings.engine.dialect.name!="postgresql":
raiseSystemExit("Run under --backend postgres (row locks are needed to reproduce).")
_seed()
unsorted_deadlocks=_run(sort_fix=False)
sorted_deadlocks=_run(sort_fix=True)
print(
f"\n=== ADRQ insert-order deadlock — {N_THREADS} threads x {N_ROUNDS} rounds, "f"{len(DOWNSTREAM_DAGS)} downstream dags ===\n"f" UNSORTED (insert in set order, forward vs reversed): {unsorted_deadlocks} deadlocks\n"f" SORTED (sort by dag_id first, the fix) : {sorted_deadlocks} deadlocks\n"
)
if__name__=="__main__":
main()
  • Yes (please specify the tool below)

Generated-by: Claude Code Opus 4.8 following the guidelines


  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

…dags
When a task updates an asset, register_asset_change queues an
AssetDagRunQueue row for each of the asset's downstream dags. The dags
come from a set, whose iteration order varies between processes, and the
rows were inserted in that order. Two tasks completing at once and
emitting to the same asset could take the per-row locks in opposite
orders and deadlock. Insert in a fixed dag_id order so the lock order is
consistent across processes; this covers the postgres, mysql, and
per-row SAVEPOINT paths.
@Andrushika
Andrushika marked this pull request as ready for review July 20, 2026 13:06
Comment threadairflow-core/src/airflow/assets/manager.py
Comment threadairflow-core/src/airflow/assets/manager.py Outdated
@Andrushika

Copy link
Copy Markdown
ContributorAuthor

After #70972, this deadlock issue does not exist anymore. Closing it.
@Dev-iL, big thanks for your review!

@Andrushika
Andrushika deleted the fix-asset-dag-run-queue-insert-deadlock branch August 7, 2026 05:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Andrushika@Dev-iL
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Fix deadlock when concurrent tasks queue the same asset's downstream dags - #70128

Closed
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock
Closed

Fix deadlock when concurrent tasks queue the same asset's downstream dags#70128
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock

Conversation

@Andrushika

Copy link
Copy Markdown
Contributor

Fix deadlock when concurrent tasks queue the same asset's downstream dags

Background

When a task produces an asset, the downstream Dags scheduled on that asset need to run. Airflow records this in the asset_dag_run_queue (ADRQ) table, one row per (asset_id, target_dag_id). So when a task succeeds and updates asset S, register_asset_change inserts one ADRQ row for every Dag scheduled on S. Those Dags are collected into a Python set (dags_to_queue), and the rows are inserted by looping over that set.

Why

A Python set has no stable order (The essence of set is a hash table). Here, the elements are DagModel objects, freshly loaded at different memory addresses in each process, so two processes can loop over the same Dags in opposite orders.

When two tasks finish at the same time and both emit to the same asset, they insert the same ADRQ rows, but maybe in opposite orders. Each insert holds a row lock until commit. So transaction A can hold row (S, dag_a) and wait for (S, dag_b), while transaction B holds (S, dag_b) and waits for (S, dag_a). That is a deadlock. Postgres aborts one side and it retries, which adds latency and error noise under high asset fan-out.

What

Insert the rows in a fixed order, sorted by dag_id, so every process takes the row locks in the same order and the cycle cannot form. A small shared helper _sorted_by_dag_id is used in all three insert paths (postgres ON CONFLICT, mysql ON DUPLICATE KEY, and the per-row SAVEPOINT fallback), because all three loop over the same set.

I was running the concurrent test for #70078 and found that this problem exists on main.
Then I reproduced it on Postgres with a script that has 8 concurrent transactions insert one asset's downstream rows in opposite orders (what two processes' set iteration can produce). Over 20 rounds it hit 658 deadlocks with the old order and 0 after sorting. I only benchmarked Postgres. MySQL uses InnoDB row locks with the same inversion, so sorting fixes it there too. The SQLite path cannot deadlock this way and is sorted only for consistency.

The script for reproduce:
from __future__ importannotationsimportthreadingimporttimefromsqlalchemyimportdeletefromsqlalchemy.dialects.postgresqlimportinsertfromsqlalchemy.excimportOperationalErrorfromairflowimportsettingsfromairflow.models.assetimportAssetActive, AssetDagRunQueue, AssetModel, DagScheduleAssetReferencefromairflow.models.dagimportDagModelfromairflow.models.dagbundleimportDagBundleModelfromairflow.utils.sessionimportcreate_sessionASSET_ID=1DOWNSTREAM_DAGS= ["dag_a", "dag_b", "dag_c", "dag_d"]
N_THREADS=8N_ROUNDS=20ROW_GAP=0.004# small pause between rows to widen the lock windowdef_seed():
withcreate_session() ass:
s.execute(delete(AssetDagRunQueue))
s.execute(delete(DagScheduleAssetReference))
s.query(AssetActive).delete()
s.query(AssetModel).delete()
s.query(DagModel).filter(DagModel.dag_id.in_(DOWNSTREAM_DAGS)).delete(synchronize_session=False)
s.merge(DagBundleModel(name="repro"))
s.flush() # bundle must exist before dags reference it (FK)asset=AssetModel(id=ASSET_ID, name="S", uri="s3://bucket/S", group="asset", extra={})
s.add_all([asset, AssetActive.for_asset(asset)])
fordinDOWNSTREAM_DAGS:
s.add(DagModel(dag_id=d, bundle_name="repro", is_stale=False, fileloc=f"{d}.py"))
s.flush()
asset.scheduled_dags= [DagScheduleAssetReference(dag_id=d) fordinDOWNSTREAM_DAGS]
def_insert_rows(order, sort_fix, counters, barrier):
dag_ids=sorted(order) ifsort_fixelseorder# the fix: always sort firstbarrier.wait()
whileTrue:
try:
withcreate_session() assession:
fordag_idindag_ids:
stmt= (
insert(AssetDagRunQueue)
.values(asset_id=ASSET_ID, target_dag_id=dag_id)
.on_conflict_do_nothing()
)
session.execute(stmt)
time.sleep(ROW_GAP)
returnexceptOperationalErrorase:
if"deadlock detected"instr(e).lower():
counters["deadlocks"] +=1continue# retry: the losing side re-runs and eventually winsraisedef_run(sort_fix):
counters= {"deadlocks": 0}
forward=DOWNSTREAM_DAGSreverse=list(reversed(DOWNSTREAM_DAGS))
for_inrange(N_ROUNDS):
withcreate_session() ass: # fresh rows each round so ON CONFLICT actually insertss.execute(delete(AssetDagRunQueue))
barrier=threading.Barrier(N_THREADS)
threads= [
threading.Thread(
target=_insert_rows,
args=(forwardifi%2==0elsereverse, sort_fix, counters, barrier),
)
foriinrange(N_THREADS)
]
fortinthreads:
t.start()
fortinthreads:
t.join()
returncounters["deadlocks"]
defmain():
ifsettings.engine.dialect.name!="postgresql":
raiseSystemExit("Run under --backend postgres (row locks are needed to reproduce).")
_seed()
unsorted_deadlocks=_run(sort_fix=False)
sorted_deadlocks=_run(sort_fix=True)
print(
f"\n=== ADRQ insert-order deadlock — {N_THREADS} threads x {N_ROUNDS} rounds, "f"{len(DOWNSTREAM_DAGS)} downstream dags ===\n"f" UNSORTED (insert in set order, forward vs reversed): {unsorted_deadlocks} deadlocks\n"f" SORTED (sort by dag_id first, the fix) : {sorted_deadlocks} deadlocks\n"
)
if__name__=="__main__":
main()
  • Yes (please specify the tool below)

Generated-by: Claude Code Opus 4.8 following the guidelines


  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

…dags
When a task updates an asset, register_asset_change queues an
AssetDagRunQueue row for each of the asset's downstream dags. The dags
come from a set, whose iteration order varies between processes, and the
rows were inserted in that order. Two tasks completing at once and
emitting to the same asset could take the per-row locks in opposite
orders and deadlock. Insert in a fixed dag_id order so the lock order is
consistent across processes; this covers the postgres, mysql, and
per-row SAVEPOINT paths.
@Andrushika
Andrushika marked this pull request as ready for review July 20, 2026 13:06
Comment threadairflow-core/src/airflow/assets/manager.py
Comment threadairflow-core/src/airflow/assets/manager.py Outdated
@Andrushika

Copy link
Copy Markdown
ContributorAuthor

After #70972, this deadlock issue does not exist anymore. Closing it.
@Dev-iL, big thanks for your review!

@Andrushika
Andrushika deleted the fix-asset-dag-run-queue-insert-deadlock branch August 7, 2026 05:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Andrushika@Dev-iL
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix deadlock when concurrent tasks queue the same asset's downstream dags - #70128

Closed
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock
Closed

Fix deadlock when concurrent tasks queue the same asset's downstream dags#70128
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock

Conversation

@Andrushika

Copy link
Copy Markdown
Contributor

Fix deadlock when concurrent tasks queue the same asset's downstream dags

Background

When a task produces an asset, the downstream Dags scheduled on that asset need to run. Airflow records this in the asset_dag_run_queue (ADRQ) table, one row per (asset_id, target_dag_id). So when a task succeeds and updates asset S, register_asset_change inserts one ADRQ row for every Dag scheduled on S. Those Dags are collected into a Python set (dags_to_queue), and the rows are inserted by looping over that set.

Why

A Python set has no stable order (The essence of set is a hash table). Here, the elements are DagModel objects, freshly loaded at different memory addresses in each process, so two processes can loop over the same Dags in opposite orders.

When two tasks finish at the same time and both emit to the same asset, they insert the same ADRQ rows, but maybe in opposite orders. Each insert holds a row lock until commit. So transaction A can hold row (S, dag_a) and wait for (S, dag_b), while transaction B holds (S, dag_b) and waits for (S, dag_a). That is a deadlock. Postgres aborts one side and it retries, which adds latency and error noise under high asset fan-out.

What

Insert the rows in a fixed order, sorted by dag_id, so every process takes the row locks in the same order and the cycle cannot form. A small shared helper _sorted_by_dag_id is used in all three insert paths (postgres ON CONFLICT, mysql ON DUPLICATE KEY, and the per-row SAVEPOINT fallback), because all three loop over the same set.

I was running the concurrent test for #70078 and found that this problem exists on main.
Then I reproduced it on Postgres with a script that has 8 concurrent transactions insert one asset's downstream rows in opposite orders (what two processes' set iteration can produce). Over 20 rounds it hit 658 deadlocks with the old order and 0 after sorting. I only benchmarked Postgres. MySQL uses InnoDB row locks with the same inversion, so sorting fixes it there too. The SQLite path cannot deadlock this way and is sorted only for consistency.

The script for reproduce:
from __future__ importannotationsimportthreadingimporttimefromsqlalchemyimportdeletefromsqlalchemy.dialects.postgresqlimportinsertfromsqlalchemy.excimportOperationalErrorfromairflowimportsettingsfromairflow.models.assetimportAssetActive, AssetDagRunQueue, AssetModel, DagScheduleAssetReferencefromairflow.models.dagimportDagModelfromairflow.models.dagbundleimportDagBundleModelfromairflow.utils.sessionimportcreate_sessionASSET_ID=1DOWNSTREAM_DAGS= ["dag_a", "dag_b", "dag_c", "dag_d"]
N_THREADS=8N_ROUNDS=20ROW_GAP=0.004# small pause between rows to widen the lock windowdef_seed():
withcreate_session() ass:
s.execute(delete(AssetDagRunQueue))
s.execute(delete(DagScheduleAssetReference))
s.query(AssetActive).delete()
s.query(AssetModel).delete()
s.query(DagModel).filter(DagModel.dag_id.in_(DOWNSTREAM_DAGS)).delete(synchronize_session=False)
s.merge(DagBundleModel(name="repro"))
s.flush() # bundle must exist before dags reference it (FK)asset=AssetModel(id=ASSET_ID, name="S", uri="s3://bucket/S", group="asset", extra={})
s.add_all([asset, AssetActive.for_asset(asset)])
fordinDOWNSTREAM_DAGS:
s.add(DagModel(dag_id=d, bundle_name="repro", is_stale=False, fileloc=f"{d}.py"))
s.flush()
asset.scheduled_dags= [DagScheduleAssetReference(dag_id=d) fordinDOWNSTREAM_DAGS]
def_insert_rows(order, sort_fix, counters, barrier):
dag_ids=sorted(order) ifsort_fixelseorder# the fix: always sort firstbarrier.wait()
whileTrue:
try:
withcreate_session() assession:
fordag_idindag_ids:
stmt= (
insert(AssetDagRunQueue)
.values(asset_id=ASSET_ID, target_dag_id=dag_id)
.on_conflict_do_nothing()
)
session.execute(stmt)
time.sleep(ROW_GAP)
returnexceptOperationalErrorase:
if"deadlock detected"instr(e).lower():
counters["deadlocks"] +=1continue# retry: the losing side re-runs and eventually winsraisedef_run(sort_fix):
counters= {"deadlocks": 0}
forward=DOWNSTREAM_DAGSreverse=list(reversed(DOWNSTREAM_DAGS))
for_inrange(N_ROUNDS):
withcreate_session() ass: # fresh rows each round so ON CONFLICT actually insertss.execute(delete(AssetDagRunQueue))
barrier=threading.Barrier(N_THREADS)
threads= [
threading.Thread(
target=_insert_rows,
args=(forwardifi%2==0elsereverse, sort_fix, counters, barrier),
)
foriinrange(N_THREADS)
]
fortinthreads:
t.start()
fortinthreads:
t.join()
returncounters["deadlocks"]
defmain():
ifsettings.engine.dialect.name!="postgresql":
raiseSystemExit("Run under --backend postgres (row locks are needed to reproduce).")
_seed()
unsorted_deadlocks=_run(sort_fix=False)
sorted_deadlocks=_run(sort_fix=True)
print(
f"\n=== ADRQ insert-order deadlock — {N_THREADS} threads x {N_ROUNDS} rounds, "f"{len(DOWNSTREAM_DAGS)} downstream dags ===\n"f" UNSORTED (insert in set order, forward vs reversed): {unsorted_deadlocks} deadlocks\n"f" SORTED (sort by dag_id first, the fix) : {sorted_deadlocks} deadlocks\n"
)
if__name__=="__main__":
main()
  • Yes (please specify the tool below)

Generated-by: Claude Code Opus 4.8 following the guidelines


  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

…dags
When a task updates an asset, register_asset_change queues an
AssetDagRunQueue row for each of the asset's downstream dags. The dags
come from a set, whose iteration order varies between processes, and the
rows were inserted in that order. Two tasks completing at once and
emitting to the same asset could take the per-row locks in opposite
orders and deadlock. Insert in a fixed dag_id order so the lock order is
consistent across processes; this covers the postgres, mysql, and
per-row SAVEPOINT paths.
@Andrushika
Andrushika marked this pull request as ready for review July 20, 2026 13:06
Comment threadairflow-core/src/airflow/assets/manager.py
Comment threadairflow-core/src/airflow/assets/manager.py Outdated
@Andrushika

Copy link
Copy Markdown
ContributorAuthor

After #70972, this deadlock issue does not exist anymore. Closing it.
@Dev-iL, big thanks for your review!

@Andrushika
Andrushika deleted the fix-asset-dag-run-queue-insert-deadlock branch August 7, 2026 05:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Andrushika@Dev-iL
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Fix deadlock when concurrent tasks queue the same asset's downstream dags - #70128

Closed
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock
Closed

Fix deadlock when concurrent tasks queue the same asset's downstream dags#70128
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock

Conversation

@Andrushika

Copy link
Copy Markdown
Contributor

Fix deadlock when concurrent tasks queue the same asset's downstream dags

Background

When a task produces an asset, the downstream Dags scheduled on that asset need to run. Airflow records this in the asset_dag_run_queue (ADRQ) table, one row per (asset_id, target_dag_id). So when a task succeeds and updates asset S, register_asset_change inserts one ADRQ row for every Dag scheduled on S. Those Dags are collected into a Python set (dags_to_queue), and the rows are inserted by looping over that set.

Why

A Python set has no stable order (The essence of set is a hash table). Here, the elements are DagModel objects, freshly loaded at different memory addresses in each process, so two processes can loop over the same Dags in opposite orders.

When two tasks finish at the same time and both emit to the same asset, they insert the same ADRQ rows, but maybe in opposite orders. Each insert holds a row lock until commit. So transaction A can hold row (S, dag_a) and wait for (S, dag_b), while transaction B holds (S, dag_b) and waits for (S, dag_a). That is a deadlock. Postgres aborts one side and it retries, which adds latency and error noise under high asset fan-out.

What

Insert the rows in a fixed order, sorted by dag_id, so every process takes the row locks in the same order and the cycle cannot form. A small shared helper _sorted_by_dag_id is used in all three insert paths (postgres ON CONFLICT, mysql ON DUPLICATE KEY, and the per-row SAVEPOINT fallback), because all three loop over the same set.

I was running the concurrent test for #70078 and found that this problem exists on main.
Then I reproduced it on Postgres with a script that has 8 concurrent transactions insert one asset's downstream rows in opposite orders (what two processes' set iteration can produce). Over 20 rounds it hit 658 deadlocks with the old order and 0 after sorting. I only benchmarked Postgres. MySQL uses InnoDB row locks with the same inversion, so sorting fixes it there too. The SQLite path cannot deadlock this way and is sorted only for consistency.

The script for reproduce:
from __future__ importannotationsimportthreadingimporttimefromsqlalchemyimportdeletefromsqlalchemy.dialects.postgresqlimportinsertfromsqlalchemy.excimportOperationalErrorfromairflowimportsettingsfromairflow.models.assetimportAssetActive, AssetDagRunQueue, AssetModel, DagScheduleAssetReferencefromairflow.models.dagimportDagModelfromairflow.models.dagbundleimportDagBundleModelfromairflow.utils.sessionimportcreate_sessionASSET_ID=1DOWNSTREAM_DAGS= ["dag_a", "dag_b", "dag_c", "dag_d"]
N_THREADS=8N_ROUNDS=20ROW_GAP=0.004# small pause between rows to widen the lock windowdef_seed():
withcreate_session() ass:
s.execute(delete(AssetDagRunQueue))
s.execute(delete(DagScheduleAssetReference))
s.query(AssetActive).delete()
s.query(AssetModel).delete()
s.query(DagModel).filter(DagModel.dag_id.in_(DOWNSTREAM_DAGS)).delete(synchronize_session=False)
s.merge(DagBundleModel(name="repro"))
s.flush() # bundle must exist before dags reference it (FK)asset=AssetModel(id=ASSET_ID, name="S", uri="s3://bucket/S", group="asset", extra={})
s.add_all([asset, AssetActive.for_asset(asset)])
fordinDOWNSTREAM_DAGS:
s.add(DagModel(dag_id=d, bundle_name="repro", is_stale=False, fileloc=f"{d}.py"))
s.flush()
asset.scheduled_dags= [DagScheduleAssetReference(dag_id=d) fordinDOWNSTREAM_DAGS]
def_insert_rows(order, sort_fix, counters, barrier):
dag_ids=sorted(order) ifsort_fixelseorder# the fix: always sort firstbarrier.wait()
whileTrue:
try:
withcreate_session() assession:
fordag_idindag_ids:
stmt= (
insert(AssetDagRunQueue)
.values(asset_id=ASSET_ID, target_dag_id=dag_id)
.on_conflict_do_nothing()
)
session.execute(stmt)
time.sleep(ROW_GAP)
returnexceptOperationalErrorase:
if"deadlock detected"instr(e).lower():
counters["deadlocks"] +=1continue# retry: the losing side re-runs and eventually winsraisedef_run(sort_fix):
counters= {"deadlocks": 0}
forward=DOWNSTREAM_DAGSreverse=list(reversed(DOWNSTREAM_DAGS))
for_inrange(N_ROUNDS):
withcreate_session() ass: # fresh rows each round so ON CONFLICT actually insertss.execute(delete(AssetDagRunQueue))
barrier=threading.Barrier(N_THREADS)
threads= [
threading.Thread(
target=_insert_rows,
args=(forwardifi%2==0elsereverse, sort_fix, counters, barrier),
)
foriinrange(N_THREADS)
]
fortinthreads:
t.start()
fortinthreads:
t.join()
returncounters["deadlocks"]
defmain():
ifsettings.engine.dialect.name!="postgresql":
raiseSystemExit("Run under --backend postgres (row locks are needed to reproduce).")
_seed()
unsorted_deadlocks=_run(sort_fix=False)
sorted_deadlocks=_run(sort_fix=True)
print(
f"\n=== ADRQ insert-order deadlock — {N_THREADS} threads x {N_ROUNDS} rounds, "f"{len(DOWNSTREAM_DAGS)} downstream dags ===\n"f" UNSORTED (insert in set order, forward vs reversed): {unsorted_deadlocks} deadlocks\n"f" SORTED (sort by dag_id first, the fix) : {sorted_deadlocks} deadlocks\n"
)
if__name__=="__main__":
main()
  • Yes (please specify the tool below)

Generated-by: Claude Code Opus 4.8 following the guidelines


  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

…dags
When a task updates an asset, register_asset_change queues an
AssetDagRunQueue row for each of the asset's downstream dags. The dags
come from a set, whose iteration order varies between processes, and the
rows were inserted in that order. Two tasks completing at once and
emitting to the same asset could take the per-row locks in opposite
orders and deadlock. Insert in a fixed dag_id order so the lock order is
consistent across processes; this covers the postgres, mysql, and
per-row SAVEPOINT paths.
@Andrushika
Andrushika marked this pull request as ready for review July 20, 2026 13:06
Comment threadairflow-core/src/airflow/assets/manager.py
Comment threadairflow-core/src/airflow/assets/manager.py Outdated
@Andrushika

Copy link
Copy Markdown
ContributorAuthor

After #70972, this deadlock issue does not exist anymore. Closing it.
@Dev-iL, big thanks for your review!

@Andrushika
Andrushika deleted the fix-asset-dag-run-queue-insert-deadlock branch August 7, 2026 05:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Andrushika@Dev-iL
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Fix deadlock when concurrent tasks queue the same asset's downstream dags - #70128

Closed
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock
Closed

Fix deadlock when concurrent tasks queue the same asset's downstream dags#70128
Andrushika wants to merge 2 commits into
apache:mainfrom
Andrushika:fix-asset-dag-run-queue-insert-deadlock

Conversation

@Andrushika

Copy link
Copy Markdown
Contributor

Fix deadlock when concurrent tasks queue the same asset's downstream dags

Background

When a task produces an asset, the downstream Dags scheduled on that asset need to run. Airflow records this in the asset_dag_run_queue (ADRQ) table, one row per (asset_id, target_dag_id). So when a task succeeds and updates asset S, register_asset_change inserts one ADRQ row for every Dag scheduled on S. Those Dags are collected into a Python set (dags_to_queue), and the rows are inserted by looping over that set.

Why

A Python set has no stable order (The essence of set is a hash table). Here, the elements are DagModel objects, freshly loaded at different memory addresses in each process, so two processes can loop over the same Dags in opposite orders.

When two tasks finish at the same time and both emit to the same asset, they insert the same ADRQ rows, but maybe in opposite orders. Each insert holds a row lock until commit. So transaction A can hold row (S, dag_a) and wait for (S, dag_b), while transaction B holds (S, dag_b) and waits for (S, dag_a). That is a deadlock. Postgres aborts one side and it retries, which adds latency and error noise under high asset fan-out.

What

Insert the rows in a fixed order, sorted by dag_id, so every process takes the row locks in the same order and the cycle cannot form. A small shared helper _sorted_by_dag_id is used in all three insert paths (postgres ON CONFLICT, mysql ON DUPLICATE KEY, and the per-row SAVEPOINT fallback), because all three loop over the same set.

I was running the concurrent test for #70078 and found that this problem exists on main.
Then I reproduced it on Postgres with a script that has 8 concurrent transactions insert one asset's downstream rows in opposite orders (what two processes' set iteration can produce). Over 20 rounds it hit 658 deadlocks with the old order and 0 after sorting. I only benchmarked Postgres. MySQL uses InnoDB row locks with the same inversion, so sorting fixes it there too. The SQLite path cannot deadlock this way and is sorted only for consistency.

The script for reproduce:
from __future__ importannotationsimportthreadingimporttimefromsqlalchemyimportdeletefromsqlalchemy.dialects.postgresqlimportinsertfromsqlalchemy.excimportOperationalErrorfromairflowimportsettingsfromairflow.models.assetimportAssetActive, AssetDagRunQueue, AssetModel, DagScheduleAssetReferencefromairflow.models.dagimportDagModelfromairflow.models.dagbundleimportDagBundleModelfromairflow.utils.sessionimportcreate_sessionASSET_ID=1DOWNSTREAM_DAGS= ["dag_a", "dag_b", "dag_c", "dag_d"]
N_THREADS=8N_ROUNDS=20ROW_GAP=0.004# small pause between rows to widen the lock windowdef_seed():
withcreate_session() ass:
s.execute(delete(AssetDagRunQueue))
s.execute(delete(DagScheduleAssetReference))
s.query(AssetActive).delete()
s.query(AssetModel).delete()
s.query(DagModel).filter(DagModel.dag_id.in_(DOWNSTREAM_DAGS)).delete(synchronize_session=False)
s.merge(DagBundleModel(name="repro"))
s.flush() # bundle must exist before dags reference it (FK)asset=AssetModel(id=ASSET_ID, name="S", uri="s3://bucket/S", group="asset", extra={})
s.add_all([asset, AssetActive.for_asset(asset)])
fordinDOWNSTREAM_DAGS:
s.add(DagModel(dag_id=d, bundle_name="repro", is_stale=False, fileloc=f"{d}.py"))
s.flush()
asset.scheduled_dags= [DagScheduleAssetReference(dag_id=d) fordinDOWNSTREAM_DAGS]
def_insert_rows(order, sort_fix, counters, barrier):
dag_ids=sorted(order) ifsort_fixelseorder# the fix: always sort firstbarrier.wait()
whileTrue:
try:
withcreate_session() assession:
fordag_idindag_ids:
stmt= (
insert(AssetDagRunQueue)
.values(asset_id=ASSET_ID, target_dag_id=dag_id)
.on_conflict_do_nothing()
)
session.execute(stmt)
time.sleep(ROW_GAP)
returnexceptOperationalErrorase:
if"deadlock detected"instr(e).lower():
counters["deadlocks"] +=1continue# retry: the losing side re-runs and eventually winsraisedef_run(sort_fix):
counters= {"deadlocks": 0}
forward=DOWNSTREAM_DAGSreverse=list(reversed(DOWNSTREAM_DAGS))
for_inrange(N_ROUNDS):
withcreate_session() ass: # fresh rows each round so ON CONFLICT actually insertss.execute(delete(AssetDagRunQueue))
barrier=threading.Barrier(N_THREADS)
threads= [
threading.Thread(
target=_insert_rows,
args=(forwardifi%2==0elsereverse, sort_fix, counters, barrier),
)
foriinrange(N_THREADS)
]
fortinthreads:
t.start()
fortinthreads:
t.join()
returncounters["deadlocks"]
defmain():
ifsettings.engine.dialect.name!="postgresql":
raiseSystemExit("Run under --backend postgres (row locks are needed to reproduce).")
_seed()
unsorted_deadlocks=_run(sort_fix=False)
sorted_deadlocks=_run(sort_fix=True)
print(
f"\n=== ADRQ insert-order deadlock — {N_THREADS} threads x {N_ROUNDS} rounds, "f"{len(DOWNSTREAM_DAGS)} downstream dags ===\n"f" UNSORTED (insert in set order, forward vs reversed): {unsorted_deadlocks} deadlocks\n"f" SORTED (sort by dag_id first, the fix) : {sorted_deadlocks} deadlocks\n"
)
if__name__=="__main__":
main()
  • Yes (please specify the tool below)

Generated-by: Claude Code Opus 4.8 following the guidelines


  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

…dags
When a task updates an asset, register_asset_change queues an
AssetDagRunQueue row for each of the asset's downstream dags. The dags
come from a set, whose iteration order varies between processes, and the
rows were inserted in that order. Two tasks completing at once and
emitting to the same asset could take the per-row locks in opposite
orders and deadlock. Insert in a fixed dag_id order so the lock order is
consistent across processes; this covers the postgres, mysql, and
per-row SAVEPOINT paths.
@Andrushika
Andrushika marked this pull request as ready for review July 20, 2026 13:06
Comment threadairflow-core/src/airflow/assets/manager.py
Comment threadairflow-core/src/airflow/assets/manager.py Outdated
@Andrushika

Copy link
Copy Markdown
ContributorAuthor

After #70972, this deadlock issue does not exist anymore. Closing it.
@Dev-iL, big thanks for your review!

@Andrushika
Andrushika deleted the fix-asset-dag-run-queue-insert-deadlock branch August 7, 2026 05:33
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Andrushika@Dev-iL