Remove "single process" restrictions on SQLite in favour of using WAL mode - #44839

Merged
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor
Dec 11, 2024
Merged

Remove "single process" restrictions on SQLite in favour of using WAL mode#44839
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor

Conversation

@ashb

@ashbashb commented Dec 11, 2024

Copy link
Copy Markdown
Member

Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.

The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support if async_mode that makes understanding the
flow complex.

Some useful docs and articles about this mode:

This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:

  • use LocalExecutor, including with more than 1 concurrent worker slot
  • have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
    that)

We execute the PRAGMA journal_mode every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.

I have tested this with breeze -b sqlite start_airflow and a kicking off a
lot of tasks concurrently.

Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've already got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!

… mode
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
def set_sqlite_pragma(dbapi_connection, connection_record):
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
cursor.execute("PRAGMA journal_mode=WAL")

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is essentially the change, everything else is removing code that isn't needed anymore!

Comment threadtests/dag_processing/test_processor.py
Comment threadairflow/dag_processing/manager.py
Comment threadairflow/dag_processing/manager.py
@ashb

ashb commented Dec 11, 2024

Copy link
Copy Markdown
MemberAuthor

Looks like I left a load of validate_database_executor_compatibility in the tests.

@ashb
ashb merged commit cb74a41 into apache:mainDec 11, 2024
@ashb
ashb deleted the remove-sync-flag-to-dagprocessor branch December 11, 2024 13:36
@potiuk

Copy link
Copy Markdown
Member

This is a fantastic improvement. And it will make "airflow standalone" finally getting really usefuil for "local experience".

ellisms pushed a commit to ellisms/airflow that referenced this pull request Dec 13, 2024
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
jedcunningham added a commit to astronomer/airflow that referenced this pull request Dec 13, 2024
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
jedcunningham added a commit that referenced this pull request Dec 14, 2024
That import was removed in #44839, but #44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in #44898 and the
"subprocess" machinery introduced in #44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in #44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:CLIarea:dev-toolsarea:Executors-coreLocalExecutor & SequentialExecutorarea:providersarea:Schedulerincluding HA (high availability) schedulerprovider:celeryprovider:cncf-kubernetesKubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@ashb@potiuk@kaxil@pierrejeambrun
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Remove "single process" restrictions on SQLite in favour of using WAL mode - #44839

Merged
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor
Dec 11, 2024
Merged

Remove "single process" restrictions on SQLite in favour of using WAL mode#44839
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor

Conversation

@ashb

@ashbashb commented Dec 11, 2024

Copy link
Copy Markdown
Member

Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.

The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support if async_mode that makes understanding the
flow complex.

Some useful docs and articles about this mode:

This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:

  • use LocalExecutor, including with more than 1 concurrent worker slot
  • have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
    that)

We execute the PRAGMA journal_mode every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.

I have tested this with breeze -b sqlite start_airflow and a kicking off a
lot of tasks concurrently.

Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've already got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!

… mode
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
def set_sqlite_pragma(dbapi_connection, connection_record):
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
cursor.execute("PRAGMA journal_mode=WAL")

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is essentially the change, everything else is removing code that isn't needed anymore!

Comment threadtests/dag_processing/test_processor.py
Comment threadairflow/dag_processing/manager.py
Comment threadairflow/dag_processing/manager.py
@ashb

ashb commented Dec 11, 2024

Copy link
Copy Markdown
MemberAuthor

Looks like I left a load of validate_database_executor_compatibility in the tests.

@ashb
ashb merged commit cb74a41 into apache:mainDec 11, 2024
@ashb
ashb deleted the remove-sync-flag-to-dagprocessor branch December 11, 2024 13:36
@potiuk

Copy link
Copy Markdown
Member

This is a fantastic improvement. And it will make "airflow standalone" finally getting really usefuil for "local experience".

ellisms pushed a commit to ellisms/airflow that referenced this pull request Dec 13, 2024
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
jedcunningham added a commit to astronomer/airflow that referenced this pull request Dec 13, 2024
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
jedcunningham added a commit that referenced this pull request Dec 14, 2024
That import was removed in #44839, but #44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in #44898 and the
"subprocess" machinery introduced in #44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in #44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:CLIarea:dev-toolsarea:Executors-coreLocalExecutor & SequentialExecutorarea:providersarea:Schedulerincluding HA (high availability) schedulerprovider:celeryprovider:cncf-kubernetesKubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@ashb@potiuk@kaxil@pierrejeambrun
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove "single process" restrictions on SQLite in favour of using WAL mode - #44839

Merged
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor
Dec 11, 2024
Merged

Remove "single process" restrictions on SQLite in favour of using WAL mode#44839
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor

Conversation

@ashb

@ashbashb commented Dec 11, 2024

Copy link
Copy Markdown
Member

Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.

The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support if async_mode that makes understanding the
flow complex.

Some useful docs and articles about this mode:

This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:

  • use LocalExecutor, including with more than 1 concurrent worker slot
  • have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
    that)

We execute the PRAGMA journal_mode every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.

I have tested this with breeze -b sqlite start_airflow and a kicking off a
lot of tasks concurrently.

Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've already got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!

… mode
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
def set_sqlite_pragma(dbapi_connection, connection_record):
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
cursor.execute("PRAGMA journal_mode=WAL")

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is essentially the change, everything else is removing code that isn't needed anymore!

Comment threadtests/dag_processing/test_processor.py
Comment threadairflow/dag_processing/manager.py
Comment threadairflow/dag_processing/manager.py
@ashb

ashb commented Dec 11, 2024

Copy link
Copy Markdown
MemberAuthor

Looks like I left a load of validate_database_executor_compatibility in the tests.

@ashb
ashb merged commit cb74a41 into apache:mainDec 11, 2024
@ashb
ashb deleted the remove-sync-flag-to-dagprocessor branch December 11, 2024 13:36
@potiuk

Copy link
Copy Markdown
Member

This is a fantastic improvement. And it will make "airflow standalone" finally getting really usefuil for "local experience".

ellisms pushed a commit to ellisms/airflow that referenced this pull request Dec 13, 2024
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
jedcunningham added a commit to astronomer/airflow that referenced this pull request Dec 13, 2024
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
jedcunningham added a commit that referenced this pull request Dec 14, 2024
That import was removed in #44839, but #44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in #44898 and the
"subprocess" machinery introduced in #44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in #44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:CLIarea:dev-toolsarea:Executors-coreLocalExecutor & SequentialExecutorarea:providersarea:Schedulerincluding HA (high availability) schedulerprovider:celeryprovider:cncf-kubernetesKubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@ashb@potiuk@kaxil@pierrejeambrun
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove "single process" restrictions on SQLite in favour of using WAL mode - #44839

Merged
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor
Dec 11, 2024
Merged

Remove "single process" restrictions on SQLite in favour of using WAL mode#44839
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor

Conversation

@ashb

@ashbashb commented Dec 11, 2024

Copy link
Copy Markdown
Member

Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.

The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support if async_mode that makes understanding the
flow complex.

Some useful docs and articles about this mode:

This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:

  • use LocalExecutor, including with more than 1 concurrent worker slot
  • have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
    that)

We execute the PRAGMA journal_mode every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.

I have tested this with breeze -b sqlite start_airflow and a kicking off a
lot of tasks concurrently.

Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've already got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!

… mode
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
def set_sqlite_pragma(dbapi_connection, connection_record):
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
cursor.execute("PRAGMA journal_mode=WAL")

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is essentially the change, everything else is removing code that isn't needed anymore!

Comment threadtests/dag_processing/test_processor.py
Comment threadairflow/dag_processing/manager.py
Comment threadairflow/dag_processing/manager.py
@ashb

ashb commented Dec 11, 2024

Copy link
Copy Markdown
MemberAuthor

Looks like I left a load of validate_database_executor_compatibility in the tests.

@ashb
ashb merged commit cb74a41 into apache:mainDec 11, 2024
@ashb
ashb deleted the remove-sync-flag-to-dagprocessor branch December 11, 2024 13:36
@potiuk

Copy link
Copy Markdown
Member

This is a fantastic improvement. And it will make "airflow standalone" finally getting really usefuil for "local experience".

ellisms pushed a commit to ellisms/airflow that referenced this pull request Dec 13, 2024
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
jedcunningham added a commit to astronomer/airflow that referenced this pull request Dec 13, 2024
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
jedcunningham added a commit that referenced this pull request Dec 14, 2024
That import was removed in #44839, but #44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in #44898 and the
"subprocess" machinery introduced in #44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in #44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:CLIarea:dev-toolsarea:Executors-coreLocalExecutor & SequentialExecutorarea:providersarea:Schedulerincluding HA (high availability) schedulerprovider:celeryprovider:cncf-kubernetesKubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@ashb@potiuk@kaxil@pierrejeambrun
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Remove "single process" restrictions on SQLite in favour of using WAL mode - #44839

Merged
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor
Dec 11, 2024
Merged

Remove "single process" restrictions on SQLite in favour of using WAL mode#44839
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor

Conversation

@ashb

@ashbashb commented Dec 11, 2024

Copy link
Copy Markdown
Member

Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.

The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support if async_mode that makes understanding the
flow complex.

Some useful docs and articles about this mode:

This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:

  • use LocalExecutor, including with more than 1 concurrent worker slot
  • have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
    that)

We execute the PRAGMA journal_mode every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.

I have tested this with breeze -b sqlite start_airflow and a kicking off a
lot of tasks concurrently.

Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've already got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!

… mode
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
def set_sqlite_pragma(dbapi_connection, connection_record):
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
cursor.execute("PRAGMA journal_mode=WAL")

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is essentially the change, everything else is removing code that isn't needed anymore!

Comment threadtests/dag_processing/test_processor.py
Comment threadairflow/dag_processing/manager.py
Comment threadairflow/dag_processing/manager.py
@ashb

ashb commented Dec 11, 2024

Copy link
Copy Markdown
MemberAuthor

Looks like I left a load of validate_database_executor_compatibility in the tests.

@ashb
ashb merged commit cb74a41 into apache:mainDec 11, 2024
@ashb
ashb deleted the remove-sync-flag-to-dagprocessor branch December 11, 2024 13:36
@potiuk

Copy link
Copy Markdown
Member

This is a fantastic improvement. And it will make "airflow standalone" finally getting really usefuil for "local experience".

ellisms pushed a commit to ellisms/airflow that referenced this pull request Dec 13, 2024
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
jedcunningham added a commit to astronomer/airflow that referenced this pull request Dec 13, 2024
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
jedcunningham added a commit that referenced this pull request Dec 14, 2024
That import was removed in #44839, but #44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in #44898 and the
"subprocess" machinery introduced in #44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in #44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:CLIarea:dev-toolsarea:Executors-coreLocalExecutor & SequentialExecutorarea:providersarea:Schedulerincluding HA (high availability) schedulerprovider:celeryprovider:cncf-kubernetesKubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@ashb@potiuk@kaxil@pierrejeambrun
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove "single process" restrictions on SQLite in favour of using WAL mode - #44839

Merged
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor
Dec 11, 2024
Merged

Remove "single process" restrictions on SQLite in favour of using WAL mode#44839
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor

Conversation

@ashb

@ashbashb commented Dec 11, 2024

Copy link
Copy Markdown
Member

Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.

The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support if async_mode that makes understanding the
flow complex.

Some useful docs and articles about this mode:

This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:

  • use LocalExecutor, including with more than 1 concurrent worker slot
  • have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
    that)

We execute the PRAGMA journal_mode every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.

I have tested this with breeze -b sqlite start_airflow and a kicking off a
lot of tasks concurrently.

Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've already got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!

… mode
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
def set_sqlite_pragma(dbapi_connection, connection_record):
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
cursor.execute("PRAGMA journal_mode=WAL")

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is essentially the change, everything else is removing code that isn't needed anymore!

Comment threadtests/dag_processing/test_processor.py
Comment threadairflow/dag_processing/manager.py
Comment threadairflow/dag_processing/manager.py
@ashb

ashb commented Dec 11, 2024

Copy link
Copy Markdown
MemberAuthor

Looks like I left a load of validate_database_executor_compatibility in the tests.

@ashb
ashb merged commit cb74a41 into apache:mainDec 11, 2024
@ashb
ashb deleted the remove-sync-flag-to-dagprocessor branch December 11, 2024 13:36
@potiuk

Copy link
Copy Markdown
Member

This is a fantastic improvement. And it will make "airflow standalone" finally getting really usefuil for "local experience".

ellisms pushed a commit to ellisms/airflow that referenced this pull request Dec 13, 2024
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
jedcunningham added a commit to astronomer/airflow that referenced this pull request Dec 13, 2024
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
jedcunningham added a commit that referenced this pull request Dec 14, 2024
That import was removed in #44839, but #44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in #44898 and the
"subprocess" machinery introduced in #44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in #44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:CLIarea:dev-toolsarea:Executors-coreLocalExecutor & SequentialExecutorarea:providersarea:Schedulerincluding HA (high availability) schedulerprovider:celeryprovider:cncf-kubernetesKubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@ashb@potiuk@kaxil@pierrejeambrun
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Remove "single process" restrictions on SQLite in favour of using WAL mode - #44839

Merged
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor
Dec 11, 2024
Merged

Remove "single process" restrictions on SQLite in favour of using WAL mode#44839
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor

Conversation

@ashb

@ashbashb commented Dec 11, 2024

Copy link
Copy Markdown
Member

Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.

The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support if async_mode that makes understanding the
flow complex.

Some useful docs and articles about this mode:

This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:

  • use LocalExecutor, including with more than 1 concurrent worker slot
  • have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
    that)

We execute the PRAGMA journal_mode every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.

I have tested this with breeze -b sqlite start_airflow and a kicking off a
lot of tasks concurrently.

Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've already got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!

… mode
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
def set_sqlite_pragma(dbapi_connection, connection_record):
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
cursor.execute("PRAGMA journal_mode=WAL")

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is essentially the change, everything else is removing code that isn't needed anymore!

Comment threadtests/dag_processing/test_processor.py
Comment threadairflow/dag_processing/manager.py
Comment threadairflow/dag_processing/manager.py
@ashb

ashb commented Dec 11, 2024

Copy link
Copy Markdown
MemberAuthor

Looks like I left a load of validate_database_executor_compatibility in the tests.

@ashb
ashb merged commit cb74a41 into apache:mainDec 11, 2024
@ashb
ashb deleted the remove-sync-flag-to-dagprocessor branch December 11, 2024 13:36
@potiuk

Copy link
Copy Markdown
Member

This is a fantastic improvement. And it will make "airflow standalone" finally getting really usefuil for "local experience".

ellisms pushed a commit to ellisms/airflow that referenced this pull request Dec 13, 2024
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
jedcunningham added a commit to astronomer/airflow that referenced this pull request Dec 13, 2024
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
jedcunningham added a commit that referenced this pull request Dec 14, 2024
That import was removed in #44839, but #44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in #44898 and the
"subprocess" machinery introduced in #44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in #44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:CLIarea:dev-toolsarea:Executors-coreLocalExecutor & SequentialExecutorarea:providersarea:Schedulerincluding HA (high availability) schedulerprovider:celeryprovider:cncf-kubernetesKubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@ashb@potiuk@kaxil@pierrejeambrun
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Remove "single process" restrictions on SQLite in favour of using WAL mode - #44839

Merged
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor
Dec 11, 2024
Merged

Remove "single process" restrictions on SQLite in favour of using WAL mode#44839
ashb merged 2 commits into
apache:mainfrom
astronomer:remove-sync-flag-to-dagprocessor

Conversation

@ashb

@ashbashb commented Dec 11, 2024

Copy link
Copy Markdown
Member

Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.

The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support if async_mode that makes understanding the
flow complex.

Some useful docs and articles about this mode:

This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:

  • use LocalExecutor, including with more than 1 concurrent worker slot
  • have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
    that)

We execute the PRAGMA journal_mode every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.

I have tested this with breeze -b sqlite start_airflow and a kicking off a
lot of tasks concurrently.

Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've already got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!

… mode
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
def set_sqlite_pragma(dbapi_connection, connection_record):
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
cursor.execute("PRAGMA journal_mode=WAL")

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is essentially the change, everything else is removing code that isn't needed anymore!

Comment threadtests/dag_processing/test_processor.py
Comment threadairflow/dag_processing/manager.py
Comment threadairflow/dag_processing/manager.py
@ashb

ashb commented Dec 11, 2024

Copy link
Copy Markdown
MemberAuthor

Looks like I left a load of validate_database_executor_compatibility in the tests.

@ashb
ashb merged commit cb74a41 into apache:mainDec 11, 2024
@ashb
ashb deleted the remove-sync-flag-to-dagprocessor branch December 11, 2024 13:36
@potiuk

Copy link
Copy Markdown
Member

This is a fantastic improvement. And it will make "airflow standalone" finally getting really usefuil for "local experience".

ellisms pushed a commit to ellisms/airflow that referenced this pull request Dec 13, 2024
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
jedcunningham added a commit to astronomer/airflow that referenced this pull request Dec 13, 2024
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
jedcunningham added a commit that referenced this pull request Dec 14, 2024
That import was removed in #44839, but #44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime objects
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 16, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 17, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 18, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit to astronomer/airflow that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMIANTE and END messages over the control socket can go.
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
ashb added a commit that referenced this pull request Dec 19, 2024
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in #44898 and the
"subprocess" machinery introduced in #44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in #44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
… mode (apache#44839)
Since 2010(!) sqlite has had a WAL, or Write-Ahead Log mode of journalling
which allos multiple concurrent readers and one writer. More than good enough
for us for "local" use.
The primary driver for this change was a realisation that it is possible and
to reduce the amount of code in complexity in DagProcessorManager before
reworking it for AIP-72 support :- we have a lot of code in the
DagProcessorManager to support `if async_mode` that makes understanding the
flow complex.
Some useful docs and articles about this mode:
- [The offical docs](https://sqlite.org/wal.html)
- [Simon Willison's TIL](https://til.simonwillison.net/sqlite/enabling-wal-mode)
- [fly.io article about scaling read concurrency](https://fly.io/blog/sqlite-internals-wal/)
This still keeps the warning against using SQLite in production, but it
greatly reduces the restrictions what combos and settings can use this. In
short, when using an SQLite db it is now possible to:
- use LocalExecutor, including with more than 1 concurrent worker slot
- have multiple DAG parsing processes (even before AIP-72/TaskSDK changes to
that)
We execute the `PRAGMA journal_mode` every time we connect, which is more
often that is strictly needed as this is one of the few modes thatis
persistent and a property of the DB file just for ease and to ensure that it
it is in the mode we want.
I have tested this with `breeze -b sqlite start_airflow` and a kicking off a
lot of tasks concurrently.
Will this be without problems? No, not entirely, but due to the
scheduler+webserver+api server process we've _already_ got the case where
multiple processes are operating on the DB file. This change just makes the
best use of that following the guidance of the SQLite project: Ensuring that
only a single process accesses the DB concurrently is not a requirement
anymore!
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
That import was removed in apache#44839, but apache#44710 wasn't up-to-date with main so
static checks there didn't fail. This simply adds it back.
got686-yandex pushed a commit to got686-yandex/airflow that referenced this pull request Jan 30, 2025
As part of Airflow 3 DAG definition files will have to use the Task SDK for
all their classes, and anything involving running user code will need to be
de-coupled from the database in the user-code process.
This change moves all of the "serialization" change up to the
DagFileProcessorManager, using the new function introduced in apache#44898 and the
"subprocess" machinery introduced in apache#44874.
**Important Note**: this change does not remove the ability for dag processes
to access the DB for Variables etc. That will come in a future change.
Some key parts of this change:
- It builds upon the WatchedSubprocess from the TaskSDK. Right now this puts a
nasty/unwanted depenednecy between the Dag Parsing code upon the TaskSDK.
This will be addressed before release (we have talked about introducing a
new "apache-airflow-base-executor" dist where this subprocess+supervisor
could live, as the "execution_time" folder in the Task SDK is more a feature
of the executor, not of the TaskSDK itself.)
- A number of classes that we need to send between processes have been
converted to Pydantic for ease of serialization.
- In order to not have to serialize everything in the subprocess and deserialize everything
in the parent Manager process, we have created a `LazyDeserializedDAG` class
that provides lazy access to much of the properties needed to create update
the DAG related DB objects, without needing to fully deserialize the entire
DAG structure.
- Classes switched to attrs based for less boilerplate in constructors.
- Internal timers convert to `time.monotonic` where possible, and `time.time`
where not, we only need second diff between two points, not datetime
objects.
- With the earlier removal of "sync mode" for SQLite in apache#44839 the need for
separate TERMINATE and END messages over the control socket can go.
---------
Co-authored-by: Jed Cunningham <66968678+jedcunningham@users.noreply.github.com>
Co-authored-by: Daniel Imberman <daniel.imberman@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:CLIarea:dev-toolsarea:Executors-coreLocalExecutor & SequentialExecutorarea:providersarea:Schedulerincluding HA (high availability) schedulerprovider:celeryprovider:cncf-kubernetesKubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@ashb@potiuk@kaxil@pierrejeambrun