Skip to content

Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor #101

Description

@ineedjet

Problem

Container renovation currently flickers: a service goes down between the old
container stopping and the new one starting. Separate from #100 (closed —
its idempotent-reconciliation proposal landed via #115: deploy/deploy.py
now runs plain docker compose pull && up -d per app, no forced restart)
— even a legitimate recreate (image actually changed) still has a
stop-then-start gap with plain docker compose up -d. This issue is
specifically about eliminating that gap, once a recreate is genuinely
warranted.

Options considered

  • Docker Swarm (rolling update, order: start-first) — rejected as
    disproportionate: requires migrating apps/networks.yml to overlay
    networks, rewriting every script's deploy command (docker stack deploy
    vs docker compose up), auditing all ~30 app compose files for
    swarm-incompatible fields (container_name, build:), and reconfiguring
    Traefik for swarmMode. A platform migration, not proportionate to "avoid
    one second of downtime."
  • Kamal — rejected: it has no concept of deploying an existing
    docker-compose.yml. It owns its own deploy manifest format
    (config/deploy.yml), proxy (kamal-proxy), and naming/labeling
    conventions, built around deploying your own app(s) you build/push
    yourself. Flightdeck's model is the opposite — every app already ships as
    a ready-made (often upstream) docker-compose.yml; adopting Kamal would
    mean translating ~30 independently-authored compose stacks into Kamal's
    format instead of just wrapping the existing up -d step.
  • docker-rollout — front
    runner. Compose-native: wraps docker compose up -d <service>, no
    rewriting of existing app compose files needed. Mechanism: scales the
    service to 2x instances, waits for the new one's healthcheck, removes the
    old one. Works with the existing Traefik Docker provider (label-based
    service discovery already handles multiple containers per router).

docker-rollout: install and constraints

Install is a single script, identical on macOS and Ubuntu — no apt/brew
package exists:

mkdir -p ~/.docker/cli-plugins
curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
chmod +x ~/.docker/cli-plugins/docker-rollout

Update: ansible/deploy.yml no longer exists (#111/#115deploy/deploy.py
now owns the entire deploy sequence, running on the CI runner, not the
target host). Installing docker-rollout would now mean adding an
idempotent bootstrap step to deploy/deploy.py's remote command sequence
(alongside the existing network/acme.json bootstrap), not an Ansible
provisioning task.

Documented caveats that constrain which apps can use it:

  • A service cannot define container_name or ports — rules out any
    app using the profile: host pattern (explicit host port mapping) per
    our own docker-compose field-ordering rules in AGENTS.md.
  • Needs a working healthcheck to know when the new instance is ready
    (falls back to a fixed wait otherwise) — common.yml's x-healthcheck is
    the base for this, but needs auditing across the catalog for actual
    adoption per app, not just presence in common.yml.
  • Not every app can safely run 2 replicas concurrently even briefly —
    e.g. anything writing to a single SQLite file on a shared volume, or
    otherwise assuming single-instance semantics. This needs a per-app
    safety judgment call, not a blanket switch.
  • Real containers get renamed with an incrementing suffix on each rollout
    (project-web-1 -> project-web-2) — worth checking this doesn't trip
    up anything relying on a stable container name (Traefik routing is
    label-based, should be fine, but worth confirming case by case).

Proposed direction

Not a blanket switch-on for the whole catalog. Roll out incrementally,
per app, starting with services where it's obviously safe (stateless,
already has a healthcheck, no host port mapping), swapping their
docker compose up -d step for docker rollout <service> in
deploy/deploy.py's per-app remote command. #100's core proposal (Compose's
own idempotent reconciliation replacing forced restarts) has since landed
via #115deploy/deploy.py runs docker compose pull && docker compose up -d --remove-orphans per app already, which is exactly the right place
to plug a per-app docker rollout swap into.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
     blocks
    (function() {
    function addCopyButtons() {
    document.querySelectorAll('pre code').forEach(function(codeBlock) {
    if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
    codeBlock.parentElement.setAttribute('data-copy-added', 'true');
    var btn = document.createElement('button');
    btn.textContent = 'Copy';
    btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
    btn.onmouseover = function() { this.style.opacity = '1'; };
    btn.onmouseout = function() { this.style.opacity = '0.7'; };
    btn.onclick = function() {
    navigator.clipboard.writeText(codeBlock.textContent).then(function() {
    btn.textContent = 'Copied!';
    setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
    });
    };
    codeBlock.parentElement.style.position = 'relative';
    codeBlock.parentElement.appendChild(btn);
    });
    }
    addCopyButtons();
    // Re-run on dynamic content
    var observer = new MutationObserver(addCopyButtons);
    observer.observe(document.body, { childList: true, subtree: true });
    })();
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor · Issue #101 · rubykatzen/flightdeck · GitHub
    Skip to content

    Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor #101

    Description

    @ineedjet

    Problem

    Container renovation currently flickers: a service goes down between the old
    container stopping and the new one starting. Separate from #100 (closed —
    its idempotent-reconciliation proposal landed via #115: deploy/deploy.py
    now runs plain docker compose pull && up -d per app, no forced restart)
    — even a legitimate recreate (image actually changed) still has a
    stop-then-start gap with plain docker compose up -d. This issue is
    specifically about eliminating that gap, once a recreate is genuinely
    warranted.

    Options considered

    • Docker Swarm (rolling update, order: start-first) — rejected as
      disproportionate: requires migrating apps/networks.yml to overlay
      networks, rewriting every script's deploy command (docker stack deploy
      vs docker compose up), auditing all ~30 app compose files for
      swarm-incompatible fields (container_name, build:), and reconfiguring
      Traefik for swarmMode. A platform migration, not proportionate to "avoid
      one second of downtime."
    • Kamal — rejected: it has no concept of deploying an existing
      docker-compose.yml. It owns its own deploy manifest format
      (config/deploy.yml), proxy (kamal-proxy), and naming/labeling
      conventions, built around deploying your own app(s) you build/push
      yourself. Flightdeck's model is the opposite — every app already ships as
      a ready-made (often upstream) docker-compose.yml; adopting Kamal would
      mean translating ~30 independently-authored compose stacks into Kamal's
      format instead of just wrapping the existing up -d step.
    • docker-rollout — front
      runner. Compose-native: wraps docker compose up -d <service>, no
      rewriting of existing app compose files needed. Mechanism: scales the
      service to 2x instances, waits for the new one's healthcheck, removes the
      old one. Works with the existing Traefik Docker provider (label-based
      service discovery already handles multiple containers per router).

    docker-rollout: install and constraints

    Install is a single script, identical on macOS and Ubuntu — no apt/brew
    package exists:

    mkdir -p ~/.docker/cli-plugins
    curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
    chmod +x ~/.docker/cli-plugins/docker-rollout

    Update: ansible/deploy.yml no longer exists (#111/#115deploy/deploy.py
    now owns the entire deploy sequence, running on the CI runner, not the
    target host). Installing docker-rollout would now mean adding an
    idempotent bootstrap step to deploy/deploy.py's remote command sequence
    (alongside the existing network/acme.json bootstrap), not an Ansible
    provisioning task.

    Documented caveats that constrain which apps can use it:

    • A service cannot define container_name or ports — rules out any
      app using the profile: host pattern (explicit host port mapping) per
      our own docker-compose field-ordering rules in AGENTS.md.
    • Needs a working healthcheck to know when the new instance is ready
      (falls back to a fixed wait otherwise) — common.yml's x-healthcheck is
      the base for this, but needs auditing across the catalog for actual
      adoption per app, not just presence in common.yml.
    • Not every app can safely run 2 replicas concurrently even briefly —
      e.g. anything writing to a single SQLite file on a shared volume, or
      otherwise assuming single-instance semantics. This needs a per-app
      safety judgment call, not a blanket switch.
    • Real containers get renamed with an incrementing suffix on each rollout
      (project-web-1 -> project-web-2) — worth checking this doesn't trip
      up anything relying on a stable container name (Traefik routing is
      label-based, should be fine, but worth confirming case by case).

    Proposed direction

    Not a blanket switch-on for the whole catalog. Roll out incrementally,
    per app, starting with services where it's obviously safe (stateless,
    already has a healthcheck, no host port mapping), swapping their
    docker compose up -d step for docker rollout <service> in
    deploy/deploy.py's per-app remote command. #100's core proposal (Compose's
    own idempotent reconciliation replacing forced restarts) has since landed
    via #115deploy/deploy.py runs docker compose pull && docker compose up -d --remove-orphans per app already, which is exactly the right place
    to plug a per-app docker rollout swap into.

    Metadata

    Metadata

    Assignees

    No one assigned

      Labels

      No labels
      No labels

      Type

      No type

      Projects

      No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor · Issue #101 · rubykatzen/flightdeck · GitHub
      Skip to content

      Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor #101

      Description

      @ineedjet

      Problem

      Container renovation currently flickers: a service goes down between the old
      container stopping and the new one starting. Separate from #100 (closed —
      its idempotent-reconciliation proposal landed via #115: deploy/deploy.py
      now runs plain docker compose pull && up -d per app, no forced restart)
      — even a legitimate recreate (image actually changed) still has a
      stop-then-start gap with plain docker compose up -d. This issue is
      specifically about eliminating that gap, once a recreate is genuinely
      warranted.

      Options considered

      • Docker Swarm (rolling update, order: start-first) — rejected as
        disproportionate: requires migrating apps/networks.yml to overlay
        networks, rewriting every script's deploy command (docker stack deploy
        vs docker compose up), auditing all ~30 app compose files for
        swarm-incompatible fields (container_name, build:), and reconfiguring
        Traefik for swarmMode. A platform migration, not proportionate to "avoid
        one second of downtime."
      • Kamal — rejected: it has no concept of deploying an existing
        docker-compose.yml. It owns its own deploy manifest format
        (config/deploy.yml), proxy (kamal-proxy), and naming/labeling
        conventions, built around deploying your own app(s) you build/push
        yourself. Flightdeck's model is the opposite — every app already ships as
        a ready-made (often upstream) docker-compose.yml; adopting Kamal would
        mean translating ~30 independently-authored compose stacks into Kamal's
        format instead of just wrapping the existing up -d step.
      • docker-rollout — front
        runner. Compose-native: wraps docker compose up -d <service>, no
        rewriting of existing app compose files needed. Mechanism: scales the
        service to 2x instances, waits for the new one's healthcheck, removes the
        old one. Works with the existing Traefik Docker provider (label-based
        service discovery already handles multiple containers per router).

      docker-rollout: install and constraints

      Install is a single script, identical on macOS and Ubuntu — no apt/brew
      package exists:

      mkdir -p ~/.docker/cli-plugins
      curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
      chmod +x ~/.docker/cli-plugins/docker-rollout

      Update: ansible/deploy.yml no longer exists (#111/#115deploy/deploy.py
      now owns the entire deploy sequence, running on the CI runner, not the
      target host). Installing docker-rollout would now mean adding an
      idempotent bootstrap step to deploy/deploy.py's remote command sequence
      (alongside the existing network/acme.json bootstrap), not an Ansible
      provisioning task.

      Documented caveats that constrain which apps can use it:

      • A service cannot define container_name or ports — rules out any
        app using the profile: host pattern (explicit host port mapping) per
        our own docker-compose field-ordering rules in AGENTS.md.
      • Needs a working healthcheck to know when the new instance is ready
        (falls back to a fixed wait otherwise) — common.yml's x-healthcheck is
        the base for this, but needs auditing across the catalog for actual
        adoption per app, not just presence in common.yml.
      • Not every app can safely run 2 replicas concurrently even briefly —
        e.g. anything writing to a single SQLite file on a shared volume, or
        otherwise assuming single-instance semantics. This needs a per-app
        safety judgment call, not a blanket switch.
      • Real containers get renamed with an incrementing suffix on each rollout
        (project-web-1 -> project-web-2) — worth checking this doesn't trip
        up anything relying on a stable container name (Traefik routing is
        label-based, should be fine, but worth confirming case by case).

      Proposed direction

      Not a blanket switch-on for the whole catalog. Roll out incrementally,
      per app, starting with services where it's obviously safe (stateless,
      already has a healthcheck, no host port mapping), swapping their
      docker compose up -d step for docker rollout <service> in
      deploy/deploy.py's per-app remote command. #100's core proposal (Compose's
      own idempotent reconciliation replacing forced restarts) has since landed
      via #115deploy/deploy.py runs docker compose pull && docker compose up -d --remove-orphans per app already, which is exactly the right place
      to plug a per-app docker rollout swap into.

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Type

        No type

        Projects

        No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor · Issue #101 · rubykatzen/flightdeck · GitHub
        Skip to content

        Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor #101

        Description

        @ineedjet

        Problem

        Container renovation currently flickers: a service goes down between the old
        container stopping and the new one starting. Separate from #100 (closed —
        its idempotent-reconciliation proposal landed via #115: deploy/deploy.py
        now runs plain docker compose pull && up -d per app, no forced restart)
        — even a legitimate recreate (image actually changed) still has a
        stop-then-start gap with plain docker compose up -d. This issue is
        specifically about eliminating that gap, once a recreate is genuinely
        warranted.

        Options considered

        • Docker Swarm (rolling update, order: start-first) — rejected as
          disproportionate: requires migrating apps/networks.yml to overlay
          networks, rewriting every script's deploy command (docker stack deploy
          vs docker compose up), auditing all ~30 app compose files for
          swarm-incompatible fields (container_name, build:), and reconfiguring
          Traefik for swarmMode. A platform migration, not proportionate to "avoid
          one second of downtime."
        • Kamal — rejected: it has no concept of deploying an existing
          docker-compose.yml. It owns its own deploy manifest format
          (config/deploy.yml), proxy (kamal-proxy), and naming/labeling
          conventions, built around deploying your own app(s) you build/push
          yourself. Flightdeck's model is the opposite — every app already ships as
          a ready-made (often upstream) docker-compose.yml; adopting Kamal would
          mean translating ~30 independently-authored compose stacks into Kamal's
          format instead of just wrapping the existing up -d step.
        • docker-rollout — front
          runner. Compose-native: wraps docker compose up -d <service>, no
          rewriting of existing app compose files needed. Mechanism: scales the
          service to 2x instances, waits for the new one's healthcheck, removes the
          old one. Works with the existing Traefik Docker provider (label-based
          service discovery already handles multiple containers per router).

        docker-rollout: install and constraints

        Install is a single script, identical on macOS and Ubuntu — no apt/brew
        package exists:

        mkdir -p ~/.docker/cli-plugins
        curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
        chmod +x ~/.docker/cli-plugins/docker-rollout

        Update: ansible/deploy.yml no longer exists (#111/#115deploy/deploy.py
        now owns the entire deploy sequence, running on the CI runner, not the
        target host). Installing docker-rollout would now mean adding an
        idempotent bootstrap step to deploy/deploy.py's remote command sequence
        (alongside the existing network/acme.json bootstrap), not an Ansible
        provisioning task.

        Documented caveats that constrain which apps can use it:

        • A service cannot define container_name or ports — rules out any
          app using the profile: host pattern (explicit host port mapping) per
          our own docker-compose field-ordering rules in AGENTS.md.
        • Needs a working healthcheck to know when the new instance is ready
          (falls back to a fixed wait otherwise) — common.yml's x-healthcheck is
          the base for this, but needs auditing across the catalog for actual
          adoption per app, not just presence in common.yml.
        • Not every app can safely run 2 replicas concurrently even briefly —
          e.g. anything writing to a single SQLite file on a shared volume, or
          otherwise assuming single-instance semantics. This needs a per-app
          safety judgment call, not a blanket switch.
        • Real containers get renamed with an incrementing suffix on each rollout
          (project-web-1 -> project-web-2) — worth checking this doesn't trip
          up anything relying on a stable container name (Traefik routing is
          label-based, should be fine, but worth confirming case by case).

        Proposed direction

        Not a blanket switch-on for the whole catalog. Roll out incrementally,
        per app, starting with services where it's obviously safe (stateless,
        already has a healthcheck, no host port mapping), swapping their
        docker compose up -d step for docker rollout <service> in
        deploy/deploy.py's per-app remote command. #100's core proposal (Compose's
        own idempotent reconciliation replacing forced restarts) has since landed
        via #115deploy/deploy.py runs docker compose pull && docker compose up -d --remove-orphans per app already, which is exactly the right place
        to plug a per-app docker rollout swap into.

        Metadata

        Metadata

        Assignees

        No one assigned

          Labels

          No labels
          No labels

          Type

          No type

          Projects

          No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor · Issue #101 · rubykatzen/flightdeck · GitHub
          Skip to content

          Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor #101

          Description

          @ineedjet

          Problem

          Container renovation currently flickers: a service goes down between the old
          container stopping and the new one starting. Separate from #100 (closed —
          its idempotent-reconciliation proposal landed via #115: deploy/deploy.py
          now runs plain docker compose pull && up -d per app, no forced restart)
          — even a legitimate recreate (image actually changed) still has a
          stop-then-start gap with plain docker compose up -d. This issue is
          specifically about eliminating that gap, once a recreate is genuinely
          warranted.

          Options considered

          • Docker Swarm (rolling update, order: start-first) — rejected as
            disproportionate: requires migrating apps/networks.yml to overlay
            networks, rewriting every script's deploy command (docker stack deploy
            vs docker compose up), auditing all ~30 app compose files for
            swarm-incompatible fields (container_name, build:), and reconfiguring
            Traefik for swarmMode. A platform migration, not proportionate to "avoid
            one second of downtime."
          • Kamal — rejected: it has no concept of deploying an existing
            docker-compose.yml. It owns its own deploy manifest format
            (config/deploy.yml), proxy (kamal-proxy), and naming/labeling
            conventions, built around deploying your own app(s) you build/push
            yourself. Flightdeck's model is the opposite — every app already ships as
            a ready-made (often upstream) docker-compose.yml; adopting Kamal would
            mean translating ~30 independently-authored compose stacks into Kamal's
            format instead of just wrapping the existing up -d step.
          • docker-rollout — front
            runner. Compose-native: wraps docker compose up -d <service>, no
            rewriting of existing app compose files needed. Mechanism: scales the
            service to 2x instances, waits for the new one's healthcheck, removes the
            old one. Works with the existing Traefik Docker provider (label-based
            service discovery already handles multiple containers per router).

          docker-rollout: install and constraints

          Install is a single script, identical on macOS and Ubuntu — no apt/brew
          package exists:

          mkdir -p ~/.docker/cli-plugins
          curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
          chmod +x ~/.docker/cli-plugins/docker-rollout

          Update: ansible/deploy.yml no longer exists (#111/#115deploy/deploy.py
          now owns the entire deploy sequence, running on the CI runner, not the
          target host). Installing docker-rollout would now mean adding an
          idempotent bootstrap step to deploy/deploy.py's remote command sequence
          (alongside the existing network/acme.json bootstrap), not an Ansible
          provisioning task.

          Documented caveats that constrain which apps can use it:

          • A service cannot define container_name or ports — rules out any
            app using the profile: host pattern (explicit host port mapping) per
            our own docker-compose field-ordering rules in AGENTS.md.
          • Needs a working healthcheck to know when the new instance is ready
            (falls back to a fixed wait otherwise) — common.yml's x-healthcheck is
            the base for this, but needs auditing across the catalog for actual
            adoption per app, not just presence in common.yml.
          • Not every app can safely run 2 replicas concurrently even briefly —
            e.g. anything writing to a single SQLite file on a shared volume, or
            otherwise assuming single-instance semantics. This needs a per-app
            safety judgment call, not a blanket switch.
          • Real containers get renamed with an incrementing suffix on each rollout
            (project-web-1 -> project-web-2) — worth checking this doesn't trip
            up anything relying on a stable container name (Traefik routing is
            label-based, should be fine, but worth confirming case by case).

          Proposed direction

          Not a blanket switch-on for the whole catalog. Roll out incrementally,
          per app, starting with services where it's obviously safe (stateless,
          already has a healthcheck, no host port mapping), swapping their
          docker compose up -d step for docker rollout <service> in
          deploy/deploy.py's per-app remote command. #100's core proposal (Compose's
          own idempotent reconciliation replacing forced restarts) has since landed
          via #115deploy/deploy.py runs docker compose pull && docker compose up -d --remove-orphans per app already, which is exactly the right place
          to plug a per-app docker rollout swap into.

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Type

            No type

            Projects

            No projects

            Milestone

            No milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor · Issue #101 · rubykatzen/flightdeck · GitHub
            Skip to content

            Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor #101

            Description

            @ineedjet

            Problem

            Container renovation currently flickers: a service goes down between the old
            container stopping and the new one starting. Separate from #100 (closed —
            its idempotent-reconciliation proposal landed via #115: deploy/deploy.py
            now runs plain docker compose pull && up -d per app, no forced restart)
            — even a legitimate recreate (image actually changed) still has a
            stop-then-start gap with plain docker compose up -d. This issue is
            specifically about eliminating that gap, once a recreate is genuinely
            warranted.

            Options considered

            • Docker Swarm (rolling update, order: start-first) — rejected as
              disproportionate: requires migrating apps/networks.yml to overlay
              networks, rewriting every script's deploy command (docker stack deploy
              vs docker compose up), auditing all ~30 app compose files for
              swarm-incompatible fields (container_name, build:), and reconfiguring
              Traefik for swarmMode. A platform migration, not proportionate to "avoid
              one second of downtime."
            • Kamal — rejected: it has no concept of deploying an existing
              docker-compose.yml. It owns its own deploy manifest format
              (config/deploy.yml), proxy (kamal-proxy), and naming/labeling
              conventions, built around deploying your own app(s) you build/push
              yourself. Flightdeck's model is the opposite — every app already ships as
              a ready-made (often upstream) docker-compose.yml; adopting Kamal would
              mean translating ~30 independently-authored compose stacks into Kamal's
              format instead of just wrapping the existing up -d step.
            • docker-rollout — front
              runner. Compose-native: wraps docker compose up -d <service>, no
              rewriting of existing app compose files needed. Mechanism: scales the
              service to 2x instances, waits for the new one's healthcheck, removes the
              old one. Works with the existing Traefik Docker provider (label-based
              service discovery already handles multiple containers per router).

            docker-rollout: install and constraints

            Install is a single script, identical on macOS and Ubuntu — no apt/brew
            package exists:

            mkdir -p ~/.docker/cli-plugins
            curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
            chmod +x ~/.docker/cli-plugins/docker-rollout

            Update: ansible/deploy.yml no longer exists (#111/#115deploy/deploy.py
            now owns the entire deploy sequence, running on the CI runner, not the
            target host). Installing docker-rollout would now mean adding an
            idempotent bootstrap step to deploy/deploy.py's remote command sequence
            (alongside the existing network/acme.json bootstrap), not an Ansible
            provisioning task.

            Documented caveats that constrain which apps can use it:

            • A service cannot define container_name or ports — rules out any
              app using the profile: host pattern (explicit host port mapping) per
              our own docker-compose field-ordering rules in AGENTS.md.
            • Needs a working healthcheck to know when the new instance is ready
              (falls back to a fixed wait otherwise) — common.yml's x-healthcheck is
              the base for this, but needs auditing across the catalog for actual
              adoption per app, not just presence in common.yml.
            • Not every app can safely run 2 replicas concurrently even briefly —
              e.g. anything writing to a single SQLite file on a shared volume, or
              otherwise assuming single-instance semantics. This needs a per-app
              safety judgment call, not a blanket switch.
            • Real containers get renamed with an incrementing suffix on each rollout
              (project-web-1 -> project-web-2) — worth checking this doesn't trip
              up anything relying on a stable container name (Traefik routing is
              label-based, should be fine, but worth confirming case by case).

            Proposed direction

            Not a blanket switch-on for the whole catalog. Roll out incrementally,
            per app, starting with services where it's obviously safe (stateless,
            already has a healthcheck, no host port mapping), swapping their
            docker compose up -d step for docker rollout <service> in
            deploy/deploy.py's per-app remote command. #100's core proposal (Compose's
            own idempotent reconciliation replacing forced restarts) has since landed
            via #115deploy/deploy.py runs docker compose pull && docker compose up -d --remove-orphans per app already, which is exactly the right place
            to plug a per-app docker rollout swap into.

            Metadata

            Metadata

            Assignees

            No one assigned

              Labels

              No labels
              No labels

              Type

              No type

              Projects

              No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor · Issue #101 · rubykatzen/flightdeck · GitHub
              Skip to content

              Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor #101

              Description

              @ineedjet

              Problem

              Container renovation currently flickers: a service goes down between the old
              container stopping and the new one starting. Separate from #100 (closed —
              its idempotent-reconciliation proposal landed via #115: deploy/deploy.py
              now runs plain docker compose pull && up -d per app, no forced restart)
              — even a legitimate recreate (image actually changed) still has a
              stop-then-start gap with plain docker compose up -d. This issue is
              specifically about eliminating that gap, once a recreate is genuinely
              warranted.

              Options considered

              • Docker Swarm (rolling update, order: start-first) — rejected as
                disproportionate: requires migrating apps/networks.yml to overlay
                networks, rewriting every script's deploy command (docker stack deploy
                vs docker compose up), auditing all ~30 app compose files for
                swarm-incompatible fields (container_name, build:), and reconfiguring
                Traefik for swarmMode. A platform migration, not proportionate to "avoid
                one second of downtime."
              • Kamal — rejected: it has no concept of deploying an existing
                docker-compose.yml. It owns its own deploy manifest format
                (config/deploy.yml), proxy (kamal-proxy), and naming/labeling
                conventions, built around deploying your own app(s) you build/push
                yourself. Flightdeck's model is the opposite — every app already ships as
                a ready-made (often upstream) docker-compose.yml; adopting Kamal would
                mean translating ~30 independently-authored compose stacks into Kamal's
                format instead of just wrapping the existing up -d step.
              • docker-rollout — front
                runner. Compose-native: wraps docker compose up -d <service>, no
                rewriting of existing app compose files needed. Mechanism: scales the
                service to 2x instances, waits for the new one's healthcheck, removes the
                old one. Works with the existing Traefik Docker provider (label-based
                service discovery already handles multiple containers per router).

              docker-rollout: install and constraints

              Install is a single script, identical on macOS and Ubuntu — no apt/brew
              package exists:

              mkdir -p ~/.docker/cli-plugins
              curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
              chmod +x ~/.docker/cli-plugins/docker-rollout

              Update: ansible/deploy.yml no longer exists (#111/#115deploy/deploy.py
              now owns the entire deploy sequence, running on the CI runner, not the
              target host). Installing docker-rollout would now mean adding an
              idempotent bootstrap step to deploy/deploy.py's remote command sequence
              (alongside the existing network/acme.json bootstrap), not an Ansible
              provisioning task.

              Documented caveats that constrain which apps can use it:

              • A service cannot define container_name or ports — rules out any
                app using the profile: host pattern (explicit host port mapping) per
                our own docker-compose field-ordering rules in AGENTS.md.
              • Needs a working healthcheck to know when the new instance is ready
                (falls back to a fixed wait otherwise) — common.yml's x-healthcheck is
                the base for this, but needs auditing across the catalog for actual
                adoption per app, not just presence in common.yml.
              • Not every app can safely run 2 replicas concurrently even briefly —
                e.g. anything writing to a single SQLite file on a shared volume, or
                otherwise assuming single-instance semantics. This needs a per-app
                safety judgment call, not a blanket switch.
              • Real containers get renamed with an incrementing suffix on each rollout
                (project-web-1 -> project-web-2) — worth checking this doesn't trip
                up anything relying on a stable container name (Traefik routing is
                label-based, should be fine, but worth confirming case by case).

              Proposed direction

              Not a blanket switch-on for the whole catalog. Roll out incrementally,
              per app, starting with services where it's obviously safe (stateless,
              already has a healthcheck, no host port mapping), swapping their
              docker compose up -d step for docker rollout <service> in
              deploy/deploy.py's per-app remote command. #100's core proposal (Compose's
              own idempotent reconciliation replacing forced restarts) has since landed
              via #115deploy/deploy.py runs docker compose pull && docker compose up -d --remove-orphans per app already, which is exactly the right place
              to plug a per-app docker rollout swap into.

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Type

                No type

                Projects

                No projects

                Milestone

                No milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions