Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 4 additions & 5 deletions docs/source/python/extending_types.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -111,15 +111,14 @@ IPC protocol::

>>> batch = pa.RecordBatch.from_arrays([arr], ["ext"])
>>> sink = pa.BufferOutputStream()
>>> writer = pa.RecordBatchStreamWriter(sink, batch.schema)
>>> writer.write_batch(batch)
>>> writer.close()
>>> with pa.RecordBatchStreamWriter(sink, batch.schema) as writer:
... writer.write_batch(batch)
>>> buf = sink.getvalue()

and then reading it back yields the proper type::

>>> reader = pa.ipc.open_stream(buf)
>>> result = reader.read_all()
>>> with pa.ipc.open_stream(buf) as reader:
... result = reader.read_all()
>>> result.column('ext').type
UuidType(extension<arrow.py_extension_type>)

Expand Down
44 changes: 23 additions & 21 deletions docs/source/python/ipc.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -65,21 +65,20 @@ this one can be created with :func:`~pyarrow.ipc.new_stream`:
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_stream(sink, batch.schema)

with pa.ipc.new_stream(sink, batch.schema) as writer:
for i in range(5):
writer.write_batch(batch)

Here we used an in-memory Arrow buffer stream, but this could have been a
socket or some other IO sink.
Here we used an in-memory Arrow buffer stream (``sink``),
but this could have been a socket or some other IO sink.

When creating the ``StreamWriter``, we pass the schema, since the schema
(column names and types) must be the same for all of the batches sent in this
particular stream. Now we can do:

.. ipython:: python

for i in range(5):
writer.write_batch(batch)
writer.close()

buf = sink.getvalue()
buf.size

Expand All@@ -89,10 +88,11 @@ convenience function ``pyarrow.ipc.open_stream``:

.. ipython:: python

reader = pa.ipc.open_stream(buf)
reader.schema

batches = [b for b in reader]
with pa.ipc.open_stream(buf) as reader:
schema = reader.schema
batches = [b for b in reader]

@amol-amol-Aug 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Took me a while to get around this one, if you wonder why this block is indented to more spaces than the other blocks, it's to work around what seemed like a bug in ipython directive.
When indented normally the second line was parsed as outside of the with block

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
In [1]: batches = [b for b in reader]

instead of

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
...: batches = [b for b in reader]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uh.


schema
len(batches)

We can check the returned batches are the same as the original input:
Expand All@@ -115,11 +115,10 @@ The :class:`~pyarrow.RecordBatchFileWriter` has the same API as
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_file(sink, batch.schema)

for i in range(10):
writer.write_batch(batch)
writer.close()

with pa.ipc.new_file(sink, batch.schema) as writer:
for i in range(10):
writer.write_batch(batch)

buf = sink.getvalue()
buf.size
Expand All@@ -131,15 +130,16 @@ operations. We can also use the :func:`~pyarrow.ipc.open_file` method to open a

.. ipython:: python

reader = pa.ipc.open_file(buf)
with pa.ipc.open_file(buf) as reader:
num_record_batches = reader.num_record_batches
b = reader.get_batch(3)

Because we have access to the entire payload, we know the number of record
batches in the file, and can read any at random:
batches in the file, and can read any at random.

.. ipython:: python

reader.num_record_batches
b = reader.get_batch(3)
num_record_batches
b.equals(batch)

Reading from Stream and File Format for pandas
Expand All@@ -151,7 +151,9 @@ DataFrame output:

.. ipython:: python

df = pa.ipc.open_file(buf).read_pandas()
with pa.ipc.open_file(buf) as reader:
df = reader.read_pandas()

df[:5]

Efficiently Writing and Reading Arrow Data
Expand Down
15 changes: 3 additions & 12 deletions docs/source/python/parquet.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -208,22 +208,13 @@ We can similarly write a Parquet file with multiple row groups by using

.. ipython:: python

writer = pq.ParquetWriter('example2.parquet', table.schema)
for i in range(3):
writer.write_table(table)
writer.close()
with pq.ParquetWriter('example2.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

pf2 = pq.ParquetFile('example2.parquet')
pf2.num_row_groups

Alternatively python ``with`` syntax can also be use:

.. ipython:: python

with pq.ParquetWriter('example3.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

Inspecting the Parquet File Metadata
------------------------------------

Expand Down
10 changes: 4 additions & 6 deletions docs/source/python/plasma.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -360,9 +360,8 @@ size of the Plasma object.
# is done to determine the size of buffer to request from the object store.
object_id = plasma.ObjectID(np.random.bytes(20))
mock_sink = pa.MockOutputStream()
stream_writer = pa.RecordBatchStreamWriter(mock_sink, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(mock_sink, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)
data_size = mock_sink.size()
buf = client.create(object_id, data_size)

Expand All@@ -372,9 +371,8 @@ The DataFrame can now be written to the buffer as follows.

# Write the PyArrow RecordBatch to Plasma
stream = pa.FixedSizeBufferWriter(buf)
stream_writer = pa.RecordBatchStreamWriter(stream, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(stream, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)

Finally, seal the finished object for use by all clients:

Expand Down
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 4 additions & 5 deletions docs/source/python/extending_types.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -111,15 +111,14 @@ IPC protocol::

>>> batch = pa.RecordBatch.from_arrays([arr], ["ext"])
>>> sink = pa.BufferOutputStream()
>>> writer = pa.RecordBatchStreamWriter(sink, batch.schema)
>>> writer.write_batch(batch)
>>> writer.close()
>>> with pa.RecordBatchStreamWriter(sink, batch.schema) as writer:
... writer.write_batch(batch)
>>> buf = sink.getvalue()

and then reading it back yields the proper type::

>>> reader = pa.ipc.open_stream(buf)
>>> result = reader.read_all()
>>> with pa.ipc.open_stream(buf) as reader:
... result = reader.read_all()
>>> result.column('ext').type
UuidType(extension<arrow.py_extension_type>)

Expand Down
44 changes: 23 additions & 21 deletions docs/source/python/ipc.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -65,21 +65,20 @@ this one can be created with :func:`~pyarrow.ipc.new_stream`:
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_stream(sink, batch.schema)

with pa.ipc.new_stream(sink, batch.schema) as writer:
for i in range(5):
writer.write_batch(batch)

Here we used an in-memory Arrow buffer stream, but this could have been a
socket or some other IO sink.
Here we used an in-memory Arrow buffer stream (``sink``),
but this could have been a socket or some other IO sink.

When creating the ``StreamWriter``, we pass the schema, since the schema
(column names and types) must be the same for all of the batches sent in this
particular stream. Now we can do:

.. ipython:: python

for i in range(5):
writer.write_batch(batch)
writer.close()

buf = sink.getvalue()
buf.size

Expand All@@ -89,10 +88,11 @@ convenience function ``pyarrow.ipc.open_stream``:

.. ipython:: python

reader = pa.ipc.open_stream(buf)
reader.schema

batches = [b for b in reader]
with pa.ipc.open_stream(buf) as reader:
schema = reader.schema
batches = [b for b in reader]

@amol-amol-Aug 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Took me a while to get around this one, if you wonder why this block is indented to more spaces than the other blocks, it's to work around what seemed like a bug in ipython directive.
When indented normally the second line was parsed as outside of the with block

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
In [1]: batches = [b for b in reader]

instead of

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
...: batches = [b for b in reader]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uh.


schema
len(batches)

We can check the returned batches are the same as the original input:
Expand All@@ -115,11 +115,10 @@ The :class:`~pyarrow.RecordBatchFileWriter` has the same API as
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_file(sink, batch.schema)

for i in range(10):
writer.write_batch(batch)
writer.close()

with pa.ipc.new_file(sink, batch.schema) as writer:
for i in range(10):
writer.write_batch(batch)

buf = sink.getvalue()
buf.size
Expand All@@ -131,15 +130,16 @@ operations. We can also use the :func:`~pyarrow.ipc.open_file` method to open a

.. ipython:: python

reader = pa.ipc.open_file(buf)
with pa.ipc.open_file(buf) as reader:
num_record_batches = reader.num_record_batches
b = reader.get_batch(3)

Because we have access to the entire payload, we know the number of record
batches in the file, and can read any at random:
batches in the file, and can read any at random.

.. ipython:: python

reader.num_record_batches
b = reader.get_batch(3)
num_record_batches
b.equals(batch)

Reading from Stream and File Format for pandas
Expand All@@ -151,7 +151,9 @@ DataFrame output:

.. ipython:: python

df = pa.ipc.open_file(buf).read_pandas()
with pa.ipc.open_file(buf) as reader:
df = reader.read_pandas()

df[:5]

Efficiently Writing and Reading Arrow Data
Expand Down
15 changes: 3 additions & 12 deletions docs/source/python/parquet.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -208,22 +208,13 @@ We can similarly write a Parquet file with multiple row groups by using

.. ipython:: python

writer = pq.ParquetWriter('example2.parquet', table.schema)
for i in range(3):
writer.write_table(table)
writer.close()
with pq.ParquetWriter('example2.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

pf2 = pq.ParquetFile('example2.parquet')
pf2.num_row_groups

Alternatively python ``with`` syntax can also be use:

.. ipython:: python

with pq.ParquetWriter('example3.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

Inspecting the Parquet File Metadata
------------------------------------

Expand Down
10 changes: 4 additions & 6 deletions docs/source/python/plasma.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -360,9 +360,8 @@ size of the Plasma object.
# is done to determine the size of buffer to request from the object store.
object_id = plasma.ObjectID(np.random.bytes(20))
mock_sink = pa.MockOutputStream()
stream_writer = pa.RecordBatchStreamWriter(mock_sink, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(mock_sink, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)
data_size = mock_sink.size()
buf = client.create(object_id, data_size)

Expand All@@ -372,9 +371,8 @@ The DataFrame can now be written to the buffer as follows.

# Write the PyArrow RecordBatch to Plasma
stream = pa.FixedSizeBufferWriter(buf)
stream_writer = pa.RecordBatchStreamWriter(stream, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(stream, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)

Finally, seal the finished object for use by all clients:

Expand Down
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 4 additions & 5 deletions docs/source/python/extending_types.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -111,15 +111,14 @@ IPC protocol::

>>> batch = pa.RecordBatch.from_arrays([arr], ["ext"])
>>> sink = pa.BufferOutputStream()
>>> writer = pa.RecordBatchStreamWriter(sink, batch.schema)
>>> writer.write_batch(batch)
>>> writer.close()
>>> with pa.RecordBatchStreamWriter(sink, batch.schema) as writer:
... writer.write_batch(batch)
>>> buf = sink.getvalue()

and then reading it back yields the proper type::

>>> reader = pa.ipc.open_stream(buf)
>>> result = reader.read_all()
>>> with pa.ipc.open_stream(buf) as reader:
... result = reader.read_all()
>>> result.column('ext').type
UuidType(extension<arrow.py_extension_type>)

Expand Down
44 changes: 23 additions & 21 deletions docs/source/python/ipc.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -65,21 +65,20 @@ this one can be created with :func:`~pyarrow.ipc.new_stream`:
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_stream(sink, batch.schema)

with pa.ipc.new_stream(sink, batch.schema) as writer:
for i in range(5):
writer.write_batch(batch)

Here we used an in-memory Arrow buffer stream, but this could have been a
socket or some other IO sink.
Here we used an in-memory Arrow buffer stream (``sink``),
but this could have been a socket or some other IO sink.

When creating the ``StreamWriter``, we pass the schema, since the schema
(column names and types) must be the same for all of the batches sent in this
particular stream. Now we can do:

.. ipython:: python

for i in range(5):
writer.write_batch(batch)
writer.close()

buf = sink.getvalue()
buf.size

Expand All@@ -89,10 +88,11 @@ convenience function ``pyarrow.ipc.open_stream``:

.. ipython:: python

reader = pa.ipc.open_stream(buf)
reader.schema

batches = [b for b in reader]
with pa.ipc.open_stream(buf) as reader:
schema = reader.schema
batches = [b for b in reader]

@amol-amol-Aug 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Took me a while to get around this one, if you wonder why this block is indented to more spaces than the other blocks, it's to work around what seemed like a bug in ipython directive.
When indented normally the second line was parsed as outside of the with block

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
In [1]: batches = [b for b in reader]

instead of

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
...: batches = [b for b in reader]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uh.


schema
len(batches)

We can check the returned batches are the same as the original input:
Expand All@@ -115,11 +115,10 @@ The :class:`~pyarrow.RecordBatchFileWriter` has the same API as
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_file(sink, batch.schema)

for i in range(10):
writer.write_batch(batch)
writer.close()

with pa.ipc.new_file(sink, batch.schema) as writer:
for i in range(10):
writer.write_batch(batch)

buf = sink.getvalue()
buf.size
Expand All@@ -131,15 +130,16 @@ operations. We can also use the :func:`~pyarrow.ipc.open_file` method to open a

.. ipython:: python

reader = pa.ipc.open_file(buf)
with pa.ipc.open_file(buf) as reader:
num_record_batches = reader.num_record_batches
b = reader.get_batch(3)

Because we have access to the entire payload, we know the number of record
batches in the file, and can read any at random:
batches in the file, and can read any at random.

.. ipython:: python

reader.num_record_batches
b = reader.get_batch(3)
num_record_batches
b.equals(batch)

Reading from Stream and File Format for pandas
Expand All@@ -151,7 +151,9 @@ DataFrame output:

.. ipython:: python

df = pa.ipc.open_file(buf).read_pandas()
with pa.ipc.open_file(buf) as reader:
df = reader.read_pandas()

df[:5]

Efficiently Writing and Reading Arrow Data
Expand Down
15 changes: 3 additions & 12 deletions docs/source/python/parquet.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -208,22 +208,13 @@ We can similarly write a Parquet file with multiple row groups by using

.. ipython:: python

writer = pq.ParquetWriter('example2.parquet', table.schema)
for i in range(3):
writer.write_table(table)
writer.close()
with pq.ParquetWriter('example2.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

pf2 = pq.ParquetFile('example2.parquet')
pf2.num_row_groups

Alternatively python ``with`` syntax can also be use:

.. ipython:: python

with pq.ParquetWriter('example3.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

Inspecting the Parquet File Metadata
------------------------------------

Expand Down
10 changes: 4 additions & 6 deletions docs/source/python/plasma.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -360,9 +360,8 @@ size of the Plasma object.
# is done to determine the size of buffer to request from the object store.
object_id = plasma.ObjectID(np.random.bytes(20))
mock_sink = pa.MockOutputStream()
stream_writer = pa.RecordBatchStreamWriter(mock_sink, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(mock_sink, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)
data_size = mock_sink.size()
buf = client.create(object_id, data_size)

Expand All@@ -372,9 +371,8 @@ The DataFrame can now be written to the buffer as follows.

# Write the PyArrow RecordBatch to Plasma
stream = pa.FixedSizeBufferWriter(buf)
stream_writer = pa.RecordBatchStreamWriter(stream, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(stream, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)

Finally, seal the finished object for use by all clients:

Expand Down
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 4 additions & 5 deletions docs/source/python/extending_types.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -111,15 +111,14 @@ IPC protocol::

>>> batch = pa.RecordBatch.from_arrays([arr], ["ext"])
>>> sink = pa.BufferOutputStream()
>>> writer = pa.RecordBatchStreamWriter(sink, batch.schema)
>>> writer.write_batch(batch)
>>> writer.close()
>>> with pa.RecordBatchStreamWriter(sink, batch.schema) as writer:
... writer.write_batch(batch)
>>> buf = sink.getvalue()

and then reading it back yields the proper type::

>>> reader = pa.ipc.open_stream(buf)
>>> result = reader.read_all()
>>> with pa.ipc.open_stream(buf) as reader:
... result = reader.read_all()
>>> result.column('ext').type
UuidType(extension<arrow.py_extension_type>)

Expand Down
44 changes: 23 additions & 21 deletions docs/source/python/ipc.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -65,21 +65,20 @@ this one can be created with :func:`~pyarrow.ipc.new_stream`:
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_stream(sink, batch.schema)

with pa.ipc.new_stream(sink, batch.schema) as writer:
for i in range(5):
writer.write_batch(batch)

Here we used an in-memory Arrow buffer stream, but this could have been a
socket or some other IO sink.
Here we used an in-memory Arrow buffer stream (``sink``),
but this could have been a socket or some other IO sink.

When creating the ``StreamWriter``, we pass the schema, since the schema
(column names and types) must be the same for all of the batches sent in this
particular stream. Now we can do:

.. ipython:: python

for i in range(5):
writer.write_batch(batch)
writer.close()

buf = sink.getvalue()
buf.size

Expand All@@ -89,10 +88,11 @@ convenience function ``pyarrow.ipc.open_stream``:

.. ipython:: python

reader = pa.ipc.open_stream(buf)
reader.schema

batches = [b for b in reader]
with pa.ipc.open_stream(buf) as reader:
schema = reader.schema
batches = [b for b in reader]

@amol-amol-Aug 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Took me a while to get around this one, if you wonder why this block is indented to more spaces than the other blocks, it's to work around what seemed like a bug in ipython directive.
When indented normally the second line was parsed as outside of the with block

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
In [1]: batches = [b for b in reader]

instead of

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
...: batches = [b for b in reader]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uh.


schema
len(batches)

We can check the returned batches are the same as the original input:
Expand All@@ -115,11 +115,10 @@ The :class:`~pyarrow.RecordBatchFileWriter` has the same API as
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_file(sink, batch.schema)

for i in range(10):
writer.write_batch(batch)
writer.close()

with pa.ipc.new_file(sink, batch.schema) as writer:
for i in range(10):
writer.write_batch(batch)

buf = sink.getvalue()
buf.size
Expand All@@ -131,15 +130,16 @@ operations. We can also use the :func:`~pyarrow.ipc.open_file` method to open a

.. ipython:: python

reader = pa.ipc.open_file(buf)
with pa.ipc.open_file(buf) as reader:
num_record_batches = reader.num_record_batches
b = reader.get_batch(3)

Because we have access to the entire payload, we know the number of record
batches in the file, and can read any at random:
batches in the file, and can read any at random.

.. ipython:: python

reader.num_record_batches
b = reader.get_batch(3)
num_record_batches
b.equals(batch)

Reading from Stream and File Format for pandas
Expand All@@ -151,7 +151,9 @@ DataFrame output:

.. ipython:: python

df = pa.ipc.open_file(buf).read_pandas()
with pa.ipc.open_file(buf) as reader:
df = reader.read_pandas()

df[:5]

Efficiently Writing and Reading Arrow Data
Expand Down
15 changes: 3 additions & 12 deletions docs/source/python/parquet.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -208,22 +208,13 @@ We can similarly write a Parquet file with multiple row groups by using

.. ipython:: python

writer = pq.ParquetWriter('example2.parquet', table.schema)
for i in range(3):
writer.write_table(table)
writer.close()
with pq.ParquetWriter('example2.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

pf2 = pq.ParquetFile('example2.parquet')
pf2.num_row_groups

Alternatively python ``with`` syntax can also be use:

.. ipython:: python

with pq.ParquetWriter('example3.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

Inspecting the Parquet File Metadata
------------------------------------

Expand Down
10 changes: 4 additions & 6 deletions docs/source/python/plasma.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -360,9 +360,8 @@ size of the Plasma object.
# is done to determine the size of buffer to request from the object store.
object_id = plasma.ObjectID(np.random.bytes(20))
mock_sink = pa.MockOutputStream()
stream_writer = pa.RecordBatchStreamWriter(mock_sink, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(mock_sink, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)
data_size = mock_sink.size()
buf = client.create(object_id, data_size)

Expand All@@ -372,9 +371,8 @@ The DataFrame can now be written to the buffer as follows.

# Write the PyArrow RecordBatch to Plasma
stream = pa.FixedSizeBufferWriter(buf)
stream_writer = pa.RecordBatchStreamWriter(stream, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(stream, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)

Finally, seal the finished object for use by all clients:

Expand Down
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 4 additions & 5 deletions docs/source/python/extending_types.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -111,15 +111,14 @@ IPC protocol::

>>> batch = pa.RecordBatch.from_arrays([arr], ["ext"])
>>> sink = pa.BufferOutputStream()
>>> writer = pa.RecordBatchStreamWriter(sink, batch.schema)
>>> writer.write_batch(batch)
>>> writer.close()
>>> with pa.RecordBatchStreamWriter(sink, batch.schema) as writer:
... writer.write_batch(batch)
>>> buf = sink.getvalue()

and then reading it back yields the proper type::

>>> reader = pa.ipc.open_stream(buf)
>>> result = reader.read_all()
>>> with pa.ipc.open_stream(buf) as reader:
... result = reader.read_all()
>>> result.column('ext').type
UuidType(extension<arrow.py_extension_type>)

Expand Down
44 changes: 23 additions & 21 deletions docs/source/python/ipc.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -65,21 +65,20 @@ this one can be created with :func:`~pyarrow.ipc.new_stream`:
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_stream(sink, batch.schema)

with pa.ipc.new_stream(sink, batch.schema) as writer:
for i in range(5):
writer.write_batch(batch)

Here we used an in-memory Arrow buffer stream, but this could have been a
socket or some other IO sink.
Here we used an in-memory Arrow buffer stream (``sink``),
but this could have been a socket or some other IO sink.

When creating the ``StreamWriter``, we pass the schema, since the schema
(column names and types) must be the same for all of the batches sent in this
particular stream. Now we can do:

.. ipython:: python

for i in range(5):
writer.write_batch(batch)
writer.close()

buf = sink.getvalue()
buf.size

Expand All@@ -89,10 +88,11 @@ convenience function ``pyarrow.ipc.open_stream``:

.. ipython:: python

reader = pa.ipc.open_stream(buf)
reader.schema

batches = [b for b in reader]
with pa.ipc.open_stream(buf) as reader:
schema = reader.schema
batches = [b for b in reader]

@amol-amol-Aug 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Took me a while to get around this one, if you wonder why this block is indented to more spaces than the other blocks, it's to work around what seemed like a bug in ipython directive.
When indented normally the second line was parsed as outside of the with block

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
In [1]: batches = [b for b in reader]

instead of

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
...: batches = [b for b in reader]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uh.


schema
len(batches)

We can check the returned batches are the same as the original input:
Expand All@@ -115,11 +115,10 @@ The :class:`~pyarrow.RecordBatchFileWriter` has the same API as
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_file(sink, batch.schema)

for i in range(10):
writer.write_batch(batch)
writer.close()

with pa.ipc.new_file(sink, batch.schema) as writer:
for i in range(10):
writer.write_batch(batch)

buf = sink.getvalue()
buf.size
Expand All@@ -131,15 +130,16 @@ operations. We can also use the :func:`~pyarrow.ipc.open_file` method to open a

.. ipython:: python

reader = pa.ipc.open_file(buf)
with pa.ipc.open_file(buf) as reader:
num_record_batches = reader.num_record_batches
b = reader.get_batch(3)

Because we have access to the entire payload, we know the number of record
batches in the file, and can read any at random:
batches in the file, and can read any at random.

.. ipython:: python

reader.num_record_batches
b = reader.get_batch(3)
num_record_batches
b.equals(batch)

Reading from Stream and File Format for pandas
Expand All@@ -151,7 +151,9 @@ DataFrame output:

.. ipython:: python

df = pa.ipc.open_file(buf).read_pandas()
with pa.ipc.open_file(buf) as reader:
df = reader.read_pandas()

df[:5]

Efficiently Writing and Reading Arrow Data
Expand Down
15 changes: 3 additions & 12 deletions docs/source/python/parquet.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -208,22 +208,13 @@ We can similarly write a Parquet file with multiple row groups by using

.. ipython:: python

writer = pq.ParquetWriter('example2.parquet', table.schema)
for i in range(3):
writer.write_table(table)
writer.close()
with pq.ParquetWriter('example2.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

pf2 = pq.ParquetFile('example2.parquet')
pf2.num_row_groups

Alternatively python ``with`` syntax can also be use:

.. ipython:: python

with pq.ParquetWriter('example3.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

Inspecting the Parquet File Metadata
------------------------------------

Expand Down
10 changes: 4 additions & 6 deletions docs/source/python/plasma.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -360,9 +360,8 @@ size of the Plasma object.
# is done to determine the size of buffer to request from the object store.
object_id = plasma.ObjectID(np.random.bytes(20))
mock_sink = pa.MockOutputStream()
stream_writer = pa.RecordBatchStreamWriter(mock_sink, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(mock_sink, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)
data_size = mock_sink.size()
buf = client.create(object_id, data_size)

Expand All@@ -372,9 +371,8 @@ The DataFrame can now be written to the buffer as follows.

# Write the PyArrow RecordBatch to Plasma
stream = pa.FixedSizeBufferWriter(buf)
stream_writer = pa.RecordBatchStreamWriter(stream, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(stream, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)

Finally, seal the finished object for use by all clients:

Expand Down
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 4 additions & 5 deletions docs/source/python/extending_types.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -111,15 +111,14 @@ IPC protocol::

>>> batch = pa.RecordBatch.from_arrays([arr], ["ext"])
>>> sink = pa.BufferOutputStream()
>>> writer = pa.RecordBatchStreamWriter(sink, batch.schema)
>>> writer.write_batch(batch)
>>> writer.close()
>>> with pa.RecordBatchStreamWriter(sink, batch.schema) as writer:
... writer.write_batch(batch)
>>> buf = sink.getvalue()

and then reading it back yields the proper type::

>>> reader = pa.ipc.open_stream(buf)
>>> result = reader.read_all()
>>> with pa.ipc.open_stream(buf) as reader:
... result = reader.read_all()
>>> result.column('ext').type
UuidType(extension<arrow.py_extension_type>)

Expand Down
44 changes: 23 additions & 21 deletions docs/source/python/ipc.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -65,21 +65,20 @@ this one can be created with :func:`~pyarrow.ipc.new_stream`:
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_stream(sink, batch.schema)

with pa.ipc.new_stream(sink, batch.schema) as writer:
for i in range(5):
writer.write_batch(batch)

Here we used an in-memory Arrow buffer stream, but this could have been a
socket or some other IO sink.
Here we used an in-memory Arrow buffer stream (``sink``),
but this could have been a socket or some other IO sink.

When creating the ``StreamWriter``, we pass the schema, since the schema
(column names and types) must be the same for all of the batches sent in this
particular stream. Now we can do:

.. ipython:: python

for i in range(5):
writer.write_batch(batch)
writer.close()

buf = sink.getvalue()
buf.size

Expand All@@ -89,10 +88,11 @@ convenience function ``pyarrow.ipc.open_stream``:

.. ipython:: python

reader = pa.ipc.open_stream(buf)
reader.schema

batches = [b for b in reader]
with pa.ipc.open_stream(buf) as reader:
schema = reader.schema
batches = [b for b in reader]

@amol-amol-Aug 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Took me a while to get around this one, if you wonder why this block is indented to more spaces than the other blocks, it's to work around what seemed like a bug in ipython directive.
When indented normally the second line was parsed as outside of the with block

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
In [1]: batches = [b for b in reader]

instead of

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
...: batches = [b for b in reader]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uh.


schema
len(batches)

We can check the returned batches are the same as the original input:
Expand All@@ -115,11 +115,10 @@ The :class:`~pyarrow.RecordBatchFileWriter` has the same API as
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_file(sink, batch.schema)

for i in range(10):
writer.write_batch(batch)
writer.close()

with pa.ipc.new_file(sink, batch.schema) as writer:
for i in range(10):
writer.write_batch(batch)

buf = sink.getvalue()
buf.size
Expand All@@ -131,15 +130,16 @@ operations. We can also use the :func:`~pyarrow.ipc.open_file` method to open a

.. ipython:: python

reader = pa.ipc.open_file(buf)
with pa.ipc.open_file(buf) as reader:
num_record_batches = reader.num_record_batches
b = reader.get_batch(3)

Because we have access to the entire payload, we know the number of record
batches in the file, and can read any at random:
batches in the file, and can read any at random.

.. ipython:: python

reader.num_record_batches
b = reader.get_batch(3)
num_record_batches
b.equals(batch)

Reading from Stream and File Format for pandas
Expand All@@ -151,7 +151,9 @@ DataFrame output:

.. ipython:: python

df = pa.ipc.open_file(buf).read_pandas()
with pa.ipc.open_file(buf) as reader:
df = reader.read_pandas()

df[:5]

Efficiently Writing and Reading Arrow Data
Expand Down
15 changes: 3 additions & 12 deletions docs/source/python/parquet.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -208,22 +208,13 @@ We can similarly write a Parquet file with multiple row groups by using

.. ipython:: python

writer = pq.ParquetWriter('example2.parquet', table.schema)
for i in range(3):
writer.write_table(table)
writer.close()
with pq.ParquetWriter('example2.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

pf2 = pq.ParquetFile('example2.parquet')
pf2.num_row_groups

Alternatively python ``with`` syntax can also be use:

.. ipython:: python

with pq.ParquetWriter('example3.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

Inspecting the Parquet File Metadata
------------------------------------

Expand Down
10 changes: 4 additions & 6 deletions docs/source/python/plasma.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -360,9 +360,8 @@ size of the Plasma object.
# is done to determine the size of buffer to request from the object store.
object_id = plasma.ObjectID(np.random.bytes(20))
mock_sink = pa.MockOutputStream()
stream_writer = pa.RecordBatchStreamWriter(mock_sink, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(mock_sink, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)
data_size = mock_sink.size()
buf = client.create(object_id, data_size)

Expand All@@ -372,9 +371,8 @@ The DataFrame can now be written to the buffer as follows.

# Write the PyArrow RecordBatch to Plasma
stream = pa.FixedSizeBufferWriter(buf)
stream_writer = pa.RecordBatchStreamWriter(stream, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(stream, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)

Finally, seal the finished object for use by all clients:

Expand Down
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 4 additions & 5 deletions docs/source/python/extending_types.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -111,15 +111,14 @@ IPC protocol::

>>> batch = pa.RecordBatch.from_arrays([arr], ["ext"])
>>> sink = pa.BufferOutputStream()
>>> writer = pa.RecordBatchStreamWriter(sink, batch.schema)
>>> writer.write_batch(batch)
>>> writer.close()
>>> with pa.RecordBatchStreamWriter(sink, batch.schema) as writer:
... writer.write_batch(batch)
>>> buf = sink.getvalue()

and then reading it back yields the proper type::

>>> reader = pa.ipc.open_stream(buf)
>>> result = reader.read_all()
>>> with pa.ipc.open_stream(buf) as reader:
... result = reader.read_all()
>>> result.column('ext').type
UuidType(extension<arrow.py_extension_type>)

Expand Down
44 changes: 23 additions & 21 deletions docs/source/python/ipc.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -65,21 +65,20 @@ this one can be created with :func:`~pyarrow.ipc.new_stream`:
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_stream(sink, batch.schema)

with pa.ipc.new_stream(sink, batch.schema) as writer:
for i in range(5):
writer.write_batch(batch)

Here we used an in-memory Arrow buffer stream, but this could have been a
socket or some other IO sink.
Here we used an in-memory Arrow buffer stream (``sink``),
but this could have been a socket or some other IO sink.

When creating the ``StreamWriter``, we pass the schema, since the schema
(column names and types) must be the same for all of the batches sent in this
particular stream. Now we can do:

.. ipython:: python

for i in range(5):
writer.write_batch(batch)
writer.close()

buf = sink.getvalue()
buf.size

Expand All@@ -89,10 +88,11 @@ convenience function ``pyarrow.ipc.open_stream``:

.. ipython:: python

reader = pa.ipc.open_stream(buf)
reader.schema

batches = [b for b in reader]
with pa.ipc.open_stream(buf) as reader:
schema = reader.schema
batches = [b for b in reader]

@amol-amol-Aug 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Took me a while to get around this one, if you wonder why this block is indented to more spaces than the other blocks, it's to work around what seemed like a bug in ipython directive.
When indented normally the second line was parsed as outside of the with block

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
In [1]: batches = [b for b in reader]

instead of

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
...: batches = [b for b in reader]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uh.


schema
len(batches)

We can check the returned batches are the same as the original input:
Expand All@@ -115,11 +115,10 @@ The :class:`~pyarrow.RecordBatchFileWriter` has the same API as
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_file(sink, batch.schema)

for i in range(10):
writer.write_batch(batch)
writer.close()

with pa.ipc.new_file(sink, batch.schema) as writer:
for i in range(10):
writer.write_batch(batch)

buf = sink.getvalue()
buf.size
Expand All@@ -131,15 +130,16 @@ operations. We can also use the :func:`~pyarrow.ipc.open_file` method to open a

.. ipython:: python

reader = pa.ipc.open_file(buf)
with pa.ipc.open_file(buf) as reader:
num_record_batches = reader.num_record_batches
b = reader.get_batch(3)

Because we have access to the entire payload, we know the number of record
batches in the file, and can read any at random:
batches in the file, and can read any at random.

.. ipython:: python

reader.num_record_batches
b = reader.get_batch(3)
num_record_batches
b.equals(batch)

Reading from Stream and File Format for pandas
Expand All@@ -151,7 +151,9 @@ DataFrame output:

.. ipython:: python

df = pa.ipc.open_file(buf).read_pandas()
with pa.ipc.open_file(buf) as reader:
df = reader.read_pandas()

df[:5]

Efficiently Writing and Reading Arrow Data
Expand Down
15 changes: 3 additions & 12 deletions docs/source/python/parquet.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -208,22 +208,13 @@ We can similarly write a Parquet file with multiple row groups by using

.. ipython:: python

writer = pq.ParquetWriter('example2.parquet', table.schema)
for i in range(3):
writer.write_table(table)
writer.close()
with pq.ParquetWriter('example2.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

pf2 = pq.ParquetFile('example2.parquet')
pf2.num_row_groups

Alternatively python ``with`` syntax can also be use:

.. ipython:: python

with pq.ParquetWriter('example3.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

Inspecting the Parquet File Metadata
------------------------------------

Expand Down
10 changes: 4 additions & 6 deletions docs/source/python/plasma.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -360,9 +360,8 @@ size of the Plasma object.
# is done to determine the size of buffer to request from the object store.
object_id = plasma.ObjectID(np.random.bytes(20))
mock_sink = pa.MockOutputStream()
stream_writer = pa.RecordBatchStreamWriter(mock_sink, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(mock_sink, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)
data_size = mock_sink.size()
buf = client.create(object_id, data_size)

Expand All@@ -372,9 +371,8 @@ The DataFrame can now be written to the buffer as follows.

# Write the PyArrow RecordBatch to Plasma
stream = pa.FixedSizeBufferWriter(buf)
stream_writer = pa.RecordBatchStreamWriter(stream, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(stream, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)

Finally, seal the finished object for use by all clients:

Expand Down
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 4 additions & 5 deletions docs/source/python/extending_types.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -111,15 +111,14 @@ IPC protocol::

>>> batch = pa.RecordBatch.from_arrays([arr], ["ext"])
>>> sink = pa.BufferOutputStream()
>>> writer = pa.RecordBatchStreamWriter(sink, batch.schema)
>>> writer.write_batch(batch)
>>> writer.close()
>>> with pa.RecordBatchStreamWriter(sink, batch.schema) as writer:
... writer.write_batch(batch)
>>> buf = sink.getvalue()

and then reading it back yields the proper type::

>>> reader = pa.ipc.open_stream(buf)
>>> result = reader.read_all()
>>> with pa.ipc.open_stream(buf) as reader:
... result = reader.read_all()
>>> result.column('ext').type
UuidType(extension<arrow.py_extension_type>)

Expand Down
44 changes: 23 additions & 21 deletions docs/source/python/ipc.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -65,21 +65,20 @@ this one can be created with :func:`~pyarrow.ipc.new_stream`:
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_stream(sink, batch.schema)

with pa.ipc.new_stream(sink, batch.schema) as writer:
for i in range(5):
writer.write_batch(batch)

Here we used an in-memory Arrow buffer stream, but this could have been a
socket or some other IO sink.
Here we used an in-memory Arrow buffer stream (``sink``),
but this could have been a socket or some other IO sink.

When creating the ``StreamWriter``, we pass the schema, since the schema
(column names and types) must be the same for all of the batches sent in this
particular stream. Now we can do:

.. ipython:: python

for i in range(5):
writer.write_batch(batch)
writer.close()

buf = sink.getvalue()
buf.size

Expand All@@ -89,10 +88,11 @@ convenience function ``pyarrow.ipc.open_stream``:

.. ipython:: python

reader = pa.ipc.open_stream(buf)
reader.schema

batches = [b for b in reader]
with pa.ipc.open_stream(buf) as reader:
schema = reader.schema
batches = [b for b in reader]

@amol-amol-Aug 18, 2021

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Took me a while to get around this one, if you wonder why this block is indented to more spaces than the other blocks, it's to work around what seemed like a bug in ipython directive.
When indented normally the second line was parsed as outside of the with block

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
In [1]: batches = [b for b in reader]

instead of

In [0]: with pa.ipc.open_stream(buf) as reader:
...: schema = reader.schema
...: batches = [b for b in reader]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Uh.


schema
len(batches)

We can check the returned batches are the same as the original input:
Expand All@@ -115,11 +115,10 @@ The :class:`~pyarrow.RecordBatchFileWriter` has the same API as
.. ipython:: python

sink = pa.BufferOutputStream()
writer = pa.ipc.new_file(sink, batch.schema)

for i in range(10):
writer.write_batch(batch)
writer.close()

with pa.ipc.new_file(sink, batch.schema) as writer:
for i in range(10):
writer.write_batch(batch)

buf = sink.getvalue()
buf.size
Expand All@@ -131,15 +130,16 @@ operations. We can also use the :func:`~pyarrow.ipc.open_file` method to open a

.. ipython:: python

reader = pa.ipc.open_file(buf)
with pa.ipc.open_file(buf) as reader:
num_record_batches = reader.num_record_batches
b = reader.get_batch(3)

Because we have access to the entire payload, we know the number of record
batches in the file, and can read any at random:
batches in the file, and can read any at random.

.. ipython:: python

reader.num_record_batches
b = reader.get_batch(3)
num_record_batches
b.equals(batch)

Reading from Stream and File Format for pandas
Expand All@@ -151,7 +151,9 @@ DataFrame output:

.. ipython:: python

df = pa.ipc.open_file(buf).read_pandas()
with pa.ipc.open_file(buf) as reader:
df = reader.read_pandas()

df[:5]

Efficiently Writing and Reading Arrow Data
Expand Down
15 changes: 3 additions & 12 deletions docs/source/python/parquet.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -208,22 +208,13 @@ We can similarly write a Parquet file with multiple row groups by using

.. ipython:: python

writer = pq.ParquetWriter('example2.parquet', table.schema)
for i in range(3):
writer.write_table(table)
writer.close()
with pq.ParquetWriter('example2.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

pf2 = pq.ParquetFile('example2.parquet')
pf2.num_row_groups

Alternatively python ``with`` syntax can also be use:

.. ipython:: python

with pq.ParquetWriter('example3.parquet', table.schema) as writer:
for i in range(3):
writer.write_table(table)

Inspecting the Parquet File Metadata
------------------------------------

Expand Down
10 changes: 4 additions & 6 deletions docs/source/python/plasma.rst
Original file line numberDiff line numberDiff line change
Expand Up@@ -360,9 +360,8 @@ size of the Plasma object.
# is done to determine the size of buffer to request from the object store.
object_id = plasma.ObjectID(np.random.bytes(20))
mock_sink = pa.MockOutputStream()
stream_writer = pa.RecordBatchStreamWriter(mock_sink, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(mock_sink, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)
data_size = mock_sink.size()
buf = client.create(object_id, data_size)

Expand All@@ -372,9 +371,8 @@ The DataFrame can now be written to the buffer as follows.

# Write the PyArrow RecordBatch to Plasma
stream = pa.FixedSizeBufferWriter(buf)
stream_writer = pa.RecordBatchStreamWriter(stream, record_batch.schema)
stream_writer.write_batch(record_batch)
stream_writer.close()
with pa.RecordBatchStreamWriter(stream, record_batch.schema) as stream_writer:
stream_writer.write_batch(record_batch)

Finally, seal the finished object for use by all clients:

Expand Down