Skip to content

Add libsvm dataset support - #32

Merged
yongtang merged 3 commits into
tensorflow:masterfrom
yupbank:add-libsvm
Dec 22, 2018
Merged

Add libsvm dataset support#32
yongtang merged 3 commits into
tensorflow:masterfrom
yupbank:add-libsvm

Conversation

@yupbank

@yupbankyupbank commented Dec 19, 2018

Copy link
Copy Markdown
Member

Address #10

add make_libsvm_dataset function, which returns a dataset contains (feature, label) per row.

def make_libsvm_dataset(file_names,
num_features,
dtype=None,
label_dtype=None,
batch_size=1,
compression_type='',
buffer_size=None,
num_parallel_parser_calls=None,
drop_final_batch=False,
prefetch_buffer_size=0):

@yupbankyupbank changed the title [WIP] Add libsvm dataset supportAdd libsvm dataset supportDec 19, 2018
Comment threadtensorflow_io/libsvm/BUILD Outdated
@yongtang

Copy link
Copy Markdown
Member

Overall looks good, though I am wondering if we could expose a class interface such as class LibSVMDataset(dataset_ops.DatasetSource)? Maybe we could add the class interface on top of the current implementation?

@yupbank

yupbank commented Dec 21, 2018

Copy link
Copy Markdown
MemberAuthor

it is hard, since we only have a parsing kernel for now, we need to implement a datasource kernel to support that basically.

if it is really worth it, i can make a second pr to port current paring kernel into datasource kernel

and the function pattern is also from tensorflow core https://github.com/tensorflow/tensorflow/blob/master/tensorflow/python/data/experimental/ops/readers.py#L311

@yongtangyongtang left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall, I think this is good. We could consider adding DatasetSource support later. LGTM

@yongtang
yongtang merged commit 726907c into tensorflow:masterDec 22, 2018
@yupbank
yupbank deleted the add-libsvm branch December 22, 2018 18:25
@yongtang

Copy link
Copy Markdown
Member

@yupbank Added a PR #38 to fix some minor issues. Please take a look.

yongtang pushed a commit that referenced this pull request Jan 27, 2022
* feat: reading from bigtable (#2)
Implements reading from bigtable in a synchronous manner.
* feat: RowRange and RowSet API.
* feat: parallel read (#4)
In this pr we make the read methods accept a row_set reading only rows specified by the user.
We also add a parallel read, that leverages the sample_row_keys method to split work among workers.
* feat: version filters (#6)
This PR adds support for Bigtable version filters.
* feat: support for other data types (#5)
* fix: linter fixes (#8)
* feat docs (#9)
* fix: building on windows (#12)
* fix: refactor bigtable package to api folder (#14)
moved bigtable to tfensorflow_io.python.api
* fix: tests hanging (#30)
changed path to bigtable emulator and cbt in tests
moved arguments' initializations to the body of the function in bigtable_ops.py
fixed interleaveFromRange of column filters when using only one column
* fix: temporarily disable macos tests (#32)
* disable tests on macos
Co-authored-by: Kajetan Boroszko <kajetan@unoperate.com>
Co-authored-by: Kajetan Boroszko <kajetan.boroszko@gmail.com>
zheolong pushed a commit to zheolong/io-1 that referenced this pull request Jul 24, 2025
* feat: reading from bigtable (tensorflow#2)
Implements reading from bigtable in a synchronous manner.
* feat: RowRange and RowSet API.
* feat: parallel read (tensorflow#4)
In this pr we make the read methods accept a row_set reading only rows specified by the user.
We also add a parallel read, that leverages the sample_row_keys method to split work among workers.
* feat: version filters (tensorflow#6)
This PR adds support for Bigtable version filters.
* feat: support for other data types (tensorflow#5)
* fix: linter fixes (tensorflow#8)
* feat docs (tensorflow#9)
* fix: building on windows (tensorflow#12)
* fix: refactor bigtable package to api folder (tensorflow#14)
moved bigtable to tfensorflow_io.python.api
* fix: tests hanging (tensorflow#30)
changed path to bigtable emulator and cbt in tests
moved arguments' initializations to the body of the function in bigtable_ops.py
fixed interleaveFromRange of column filters when using only one column
* fix: temporarily disable macos tests (tensorflow#32)
* disable tests on macos
Co-authored-by: Kajetan Boroszko <kajetan@unoperate.com>
Co-authored-by: Kajetan Boroszko <kajetan.boroszko@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@yupbank@yongtang