[C++][Dataset] Ability to read/subset a ParquetFileFragment with given set of row group ids #18298

Description

@asfimport

From discussion at dask/dask#6534 (comment) (dask using the dataset API in their parquet reader), it might be useful to somehow "subset" or read a subset of a ParquetFileFragment for a specific set of row group ids.

Use cases:

  • Read only a set of row groups ids (this is similar as ParquetFile.read_row_groups), eg because you want to control the size of the resulting table by reading subsets of row groups

  • Get a ParquetFileFragment with a subset of row groups (eg based on a filter) to then eg get the statistics of only those row groups

    The first case could for example be solved by adding a row_groups keyword to ParquetFileFragment.to_table (but, this is then a keyword specific to the parquet format, and we should then probably also add it to scan et al).

    The second case is something you can in principle do yourself manually by recreating a fragment with fragment.format.make_fragment(fragment.path, ..., row_groups=[...]). However, this is a) a bit cumbersome and b) statistics might need to be parsed again?
    The statistics of a set of filtered row groups could also be obtained by using split_by_row_group(filter) (and then get the statistics of each of the fragments), but if you then want a single fragment, you need to recreate a fragment with the obtained row group ids.

    So one idea I have now (but mostly brainstorming here). Would it be useful to have a method to create a "subsetted" ParquetFileFragment, either based on a list of row group ids (fragment.subset(row_groups=[...]) or either based on a filter (fragment.subset(filter=...), which would be equivalent as split_by_row_group+recombining into a single fragment) ?

    cc @bkietz@rjzamora

Reporter: Joris Van den Bossche / @jorisvandenbossche
Assignee: Joris Van den Bossche / @jorisvandenbossche

PRs and other links:

Note: This issue was originally created as ARROW-10100. Please see the migration documentation for further details.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
     blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    
    Skip to content

    [C++][Dataset] Ability to read/subset a ParquetFileFragment with given set of row group ids #18298

    Description

    @asfimport

    From discussion at dask/dask#6534 (comment) (dask using the dataset API in their parquet reader), it might be useful to somehow "subset" or read a subset of a ParquetFileFragment for a specific set of row group ids.

    Use cases:

    • Read only a set of row groups ids (this is similar as ParquetFile.read_row_groups), eg because you want to control the size of the resulting table by reading subsets of row groups

    • Get a ParquetFileFragment with a subset of row groups (eg based on a filter) to then eg get the statistics of only those row groups

      The first case could for example be solved by adding a row_groups keyword to ParquetFileFragment.to_table (but, this is then a keyword specific to the parquet format, and we should then probably also add it to scan et al).

      The second case is something you can in principle do yourself manually by recreating a fragment with fragment.format.make_fragment(fragment.path, ..., row_groups=[...]). However, this is a) a bit cumbersome and b) statistics might need to be parsed again?
      The statistics of a set of filtered row groups could also be obtained by using split_by_row_group(filter) (and then get the statistics of each of the fragments), but if you then want a single fragment, you need to recreate a fragment with the obtained row group ids.

      So one idea I have now (but mostly brainstorming here). Would it be useful to have a method to create a "subsetted" ParquetFileFragment, either based on a list of row group ids (fragment.subset(row_groups=[...]) or either based on a filter (fragment.subset(filter=...), which would be equivalent as split_by_row_group+recombining into a single fragment) ?

      cc @bkietz@rjzamora

    Reporter: Joris Van den Bossche / @jorisvandenbossche
    Assignee: Joris Van den Bossche / @jorisvandenbossche

    PRs and other links:

    Note: This issue was originally created as ARROW-10100. Please see the migration documentation for further details.

    Activity

    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
      Skip to content

      [C++][Dataset] Ability to read/subset a ParquetFileFragment with given set of row group ids #18298

      Description

      @asfimport

      From discussion at dask/dask#6534 (comment) (dask using the dataset API in their parquet reader), it might be useful to somehow "subset" or read a subset of a ParquetFileFragment for a specific set of row group ids.

      Use cases:

      • Read only a set of row groups ids (this is similar as ParquetFile.read_row_groups), eg because you want to control the size of the resulting table by reading subsets of row groups

      • Get a ParquetFileFragment with a subset of row groups (eg based on a filter) to then eg get the statistics of only those row groups

        The first case could for example be solved by adding a row_groups keyword to ParquetFileFragment.to_table (but, this is then a keyword specific to the parquet format, and we should then probably also add it to scan et al).

        The second case is something you can in principle do yourself manually by recreating a fragment with fragment.format.make_fragment(fragment.path, ..., row_groups=[...]). However, this is a) a bit cumbersome and b) statistics might need to be parsed again?
        The statistics of a set of filtered row groups could also be obtained by using split_by_row_group(filter) (and then get the statistics of each of the fragments), but if you then want a single fragment, you need to recreate a fragment with the obtained row group ids.

        So one idea I have now (but mostly brainstorming here). Would it be useful to have a method to create a "subsetted" ParquetFileFragment, either based on a list of row group ids (fragment.subset(row_groups=[...]) or either based on a filter (fragment.subset(filter=...), which would be equivalent as split_by_row_group+recombining into a single fragment) ?

        cc @bkietz@rjzamora

      Reporter: Joris Van den Bossche / @jorisvandenbossche
      Assignee: Joris Van den Bossche / @jorisvandenbossche

      PRs and other links:

      Note: This issue was originally created as ARROW-10100. Please see the migration documentation for further details.

      Activity

      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

      Metadata

      Metadata

      Type

      No type

      Projects

      No projects

        Milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
        Skip to content

        [C++][Dataset] Ability to read/subset a ParquetFileFragment with given set of row group ids #18298

        Description

        @asfimport

        From discussion at dask/dask#6534 (comment) (dask using the dataset API in their parquet reader), it might be useful to somehow "subset" or read a subset of a ParquetFileFragment for a specific set of row group ids.

        Use cases:

        • Read only a set of row groups ids (this is similar as ParquetFile.read_row_groups), eg because you want to control the size of the resulting table by reading subsets of row groups

        • Get a ParquetFileFragment with a subset of row groups (eg based on a filter) to then eg get the statistics of only those row groups

          The first case could for example be solved by adding a row_groups keyword to ParquetFileFragment.to_table (but, this is then a keyword specific to the parquet format, and we should then probably also add it to scan et al).

          The second case is something you can in principle do yourself manually by recreating a fragment with fragment.format.make_fragment(fragment.path, ..., row_groups=[...]). However, this is a) a bit cumbersome and b) statistics might need to be parsed again?
          The statistics of a set of filtered row groups could also be obtained by using split_by_row_group(filter) (and then get the statistics of each of the fragments), but if you then want a single fragment, you need to recreate a fragment with the obtained row group ids.

          So one idea I have now (but mostly brainstorming here). Would it be useful to have a method to create a "subsetted" ParquetFileFragment, either based on a list of row group ids (fragment.subset(row_groups=[...]) or either based on a filter (fragment.subset(filter=...), which would be equivalent as split_by_row_group+recombining into a single fragment) ?

          cc @bkietz@rjzamora

        Reporter: Joris Van den Bossche / @jorisvandenbossche
        Assignee: Joris Van den Bossche / @jorisvandenbossche

        PRs and other links:

        Note: This issue was originally created as ARROW-10100. Please see the migration documentation for further details.

        Activity

        Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

        Metadata

        Metadata

        Type

        No type

        Projects

        No projects

          Milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
          Skip to content

          [C++][Dataset] Ability to read/subset a ParquetFileFragment with given set of row group ids #18298

          Description

          @asfimport

          From discussion at dask/dask#6534 (comment) (dask using the dataset API in their parquet reader), it might be useful to somehow "subset" or read a subset of a ParquetFileFragment for a specific set of row group ids.

          Use cases:

          • Read only a set of row groups ids (this is similar as ParquetFile.read_row_groups), eg because you want to control the size of the resulting table by reading subsets of row groups

          • Get a ParquetFileFragment with a subset of row groups (eg based on a filter) to then eg get the statistics of only those row groups

            The first case could for example be solved by adding a row_groups keyword to ParquetFileFragment.to_table (but, this is then a keyword specific to the parquet format, and we should then probably also add it to scan et al).

            The second case is something you can in principle do yourself manually by recreating a fragment with fragment.format.make_fragment(fragment.path, ..., row_groups=[...]). However, this is a) a bit cumbersome and b) statistics might need to be parsed again?
            The statistics of a set of filtered row groups could also be obtained by using split_by_row_group(filter) (and then get the statistics of each of the fragments), but if you then want a single fragment, you need to recreate a fragment with the obtained row group ids.

            So one idea I have now (but mostly brainstorming here). Would it be useful to have a method to create a "subsetted" ParquetFileFragment, either based on a list of row group ids (fragment.subset(row_groups=[...]) or either based on a filter (fragment.subset(filter=...), which would be equivalent as split_by_row_group+recombining into a single fragment) ?

            cc @bkietz@rjzamora

          Reporter: Joris Van den Bossche / @jorisvandenbossche
          Assignee: Joris Van den Bossche / @jorisvandenbossche

          PRs and other links:

          Note: This issue was originally created as ARROW-10100. Please see the migration documentation for further details.

          Activity

          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

          Metadata

          Metadata

          Type

          No type

          Projects

          No projects

            Milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
            Skip to content

            [C++][Dataset] Ability to read/subset a ParquetFileFragment with given set of row group ids #18298

            Description

            @asfimport

            From discussion at dask/dask#6534 (comment) (dask using the dataset API in their parquet reader), it might be useful to somehow "subset" or read a subset of a ParquetFileFragment for a specific set of row group ids.

            Use cases:

            • Read only a set of row groups ids (this is similar as ParquetFile.read_row_groups), eg because you want to control the size of the resulting table by reading subsets of row groups

            • Get a ParquetFileFragment with a subset of row groups (eg based on a filter) to then eg get the statistics of only those row groups

              The first case could for example be solved by adding a row_groups keyword to ParquetFileFragment.to_table (but, this is then a keyword specific to the parquet format, and we should then probably also add it to scan et al).

              The second case is something you can in principle do yourself manually by recreating a fragment with fragment.format.make_fragment(fragment.path, ..., row_groups=[...]). However, this is a) a bit cumbersome and b) statistics might need to be parsed again?
              The statistics of a set of filtered row groups could also be obtained by using split_by_row_group(filter) (and then get the statistics of each of the fragments), but if you then want a single fragment, you need to recreate a fragment with the obtained row group ids.

              So one idea I have now (but mostly brainstorming here). Would it be useful to have a method to create a "subsetted" ParquetFileFragment, either based on a list of row group ids (fragment.subset(row_groups=[...]) or either based on a filter (fragment.subset(filter=...), which would be equivalent as split_by_row_group+recombining into a single fragment) ?

              cc @bkietz@rjzamora

            Reporter: Joris Van den Bossche / @jorisvandenbossche
            Assignee: Joris Van den Bossche / @jorisvandenbossche

            PRs and other links:

            Note: This issue was originally created as ARROW-10100. Please see the migration documentation for further details.

            Activity

            Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

            Metadata

            Metadata

            Type

            No type

            Projects

            No projects

              Milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              [C++][Dataset] Ability to read/subset a ParquetFileFragment with given set of row group ids #18298

              Description

              @asfimport

              From discussion at dask/dask#6534 (comment) (dask using the dataset API in their parquet reader), it might be useful to somehow "subset" or read a subset of a ParquetFileFragment for a specific set of row group ids.

              Use cases:

              • Read only a set of row groups ids (this is similar as ParquetFile.read_row_groups), eg because you want to control the size of the resulting table by reading subsets of row groups

              • Get a ParquetFileFragment with a subset of row groups (eg based on a filter) to then eg get the statistics of only those row groups

                The first case could for example be solved by adding a row_groups keyword to ParquetFileFragment.to_table (but, this is then a keyword specific to the parquet format, and we should then probably also add it to scan et al).

                The second case is something you can in principle do yourself manually by recreating a fragment with fragment.format.make_fragment(fragment.path, ..., row_groups=[...]). However, this is a) a bit cumbersome and b) statistics might need to be parsed again?
                The statistics of a set of filtered row groups could also be obtained by using split_by_row_group(filter) (and then get the statistics of each of the fragments), but if you then want a single fragment, you need to recreate a fragment with the obtained row group ids.

                So one idea I have now (but mostly brainstorming here). Would it be useful to have a method to create a "subsetted" ParquetFileFragment, either based on a list of row group ids (fragment.subset(row_groups=[...]) or either based on a filter (fragment.subset(filter=...), which would be equivalent as split_by_row_group+recombining into a single fragment) ?

                cc @bkietz@rjzamora

              Reporter: Joris Van den Bossche / @jorisvandenbossche
              Assignee: Joris Van den Bossche / @jorisvandenbossche

              PRs and other links:

              Note: This issue was originally created as ARROW-10100. Please see the migration documentation for further details.

              Activity

              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

              Metadata

              Metadata

              Type

              No type

              Projects

              No projects

                Milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions

                , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                Skip to content

                [C++][Dataset] Ability to read/subset a ParquetFileFragment with given set of row group ids #18298

                Description

                @asfimport

                From discussion at dask/dask#6534 (comment) (dask using the dataset API in their parquet reader), it might be useful to somehow "subset" or read a subset of a ParquetFileFragment for a specific set of row group ids.

                Use cases:

                • Read only a set of row groups ids (this is similar as ParquetFile.read_row_groups), eg because you want to control the size of the resulting table by reading subsets of row groups

                • Get a ParquetFileFragment with a subset of row groups (eg based on a filter) to then eg get the statistics of only those row groups

                  The first case could for example be solved by adding a row_groups keyword to ParquetFileFragment.to_table (but, this is then a keyword specific to the parquet format, and we should then probably also add it to scan et al).

                  The second case is something you can in principle do yourself manually by recreating a fragment with fragment.format.make_fragment(fragment.path, ..., row_groups=[...]). However, this is a) a bit cumbersome and b) statistics might need to be parsed again?
                  The statistics of a set of filtered row groups could also be obtained by using split_by_row_group(filter) (and then get the statistics of each of the fragments), but if you then want a single fragment, you need to recreate a fragment with the obtained row group ids.

                  So one idea I have now (but mostly brainstorming here). Would it be useful to have a method to create a "subsetted" ParquetFileFragment, either based on a list of row group ids (fragment.subset(row_groups=[...]) or either based on a filter (fragment.subset(filter=...), which would be equivalent as split_by_row_group+recombining into a single fragment) ?

                  cc @bkietz@rjzamora

                Reporter: Joris Van den Bossche / @jorisvandenbossche
                Assignee: Joris Van den Bossche / @jorisvandenbossche

                PRs and other links:

                Note: This issue was originally created as ARROW-10100. Please see the migration documentation for further details.

                Activity

                Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                Metadata

                Metadata

                Type

                No type

                Projects

                No projects

                  Milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions