Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TwiGet

TwiGet is a python package for the management of the queries on filtered stream of the Twitter API, and the collection of tweets from it.

It can be used as a command line tool (twiget-cli) or as a python class (TwiGet).

Installation

> pip install twiget

The command installs the package and also makes the twiget-cli command available.

Command line tool: twiget-cli

TwiGet implements a command line interface that can be started with the command:

> twiget-cli

When launched without arguments the program searches for a .twiget.conf file in the HOME directory (the directory pointed by the $HOME or %userprofile% environment variable). The file must contain in the first line the bearer token that allows the program to access the Twitter API.

Alternatively, the name of the file from which to obtain the bearer token can be given as argument when starting the program:

> twiget-cli -b path_to_file/with_token.txt

NOTE: store the bearer token in a file with minimum access permissions. Never share it. Revoke any tokens that may have been made public.

Another optional argument is the path where to save collected tweets. By default, a data subdirectory is created in the current working directory.

> twiget-cli -s ./save_dir

prompt

When started, twiget-cli shown the available commands, and the queries currently registered for the given bearer token (queries are permanently stored on Twitter's servers).

TwiGet 0.1.1
Available commands (type help <command> for details):
create, delete, exit, help, list, refresh, save_to, size, start, stop
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=1405490304970434817	query="bts"	tag="bts"

The command prompt tells if twiget-cli is currently collecting tweets, the number of collected tweets, and the save path.

[not collecting (0 since last start), save path "data"]>

When collecting tweets, the prompt is automatically refreshed every time a given number of tweets is collected (see the refresh command).

Commands

create

Format:

> create <tag> <query>

Creates a filtering rule, associated to a given tag name.
Collected tweets are saved in json format in a file named <tag>.json, in the given save path. Tag name is the first argument of the command and cannot contain spaces.
Any word after the tag defines the query. Info on how to define rules.

Example:

[not collecting (0 since last start), save path "data"]>create usa jow biden
Tweets matching the query "jow biden" will be saved in the file data/usa.json
ID=1395720345987340524

list

Format:

> list

Lists the queries, their ID and their tag, currently registered for the filtered stream.

Example:

[not collecting (0 since last start), save path "data"]> list
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=13905490304970434817	query="bts"	tag="bts"
ID=1395720345987340524	query="joe biden"	tag="usa"

delete

Format:

> delete <ID>

Deletes a query, given its ID.

Example:

[not collecting (0 since last start), save path "data"]> delete 1385892384573355842

start

Format:

> start

Starts a background process that collects tweets from the filtered stream and puts them in json files, according to the tag they are associated to.

Collection continues until a stop or a exit command is entered. To let TwiGet collect data for longer periods of time, I suggest to use TwiGet within a virtual terminal session, using, e.g., screen or tmux.

Note: create and delete command can be issued also when collecting tweets. The collection process is updated immediately.

Example:

[not collecting (0 since last start), save path "data"]> start
[collecting (0 since last start), save path "data"]>

stop

Format:

> stop

Stop data collections.

Example:

[collecting (3000 since last start), save path "data"]> stop
[not collecting (3152 since last start), save path "data"]> 

save_to

Format:

> save_to <path>

Sets the path where json files are saved.

_Note: changing path while collecting tweets will immediately create new json file in the new path, leaving all tweets collected until that moment in the old path.

Example:

[not collecting (0 since last start), save path "data"]> save_to ../my_project
[not collecting (0 since last start), save path "../my_project"]> 

size

Format:

> size <size>

Sets the maximum size in bytes of json files. When a json file reaches this size, a new file with an incremented index (e.g., tag_0.json, tag_1.json, tag_2.json...) is created.

Example:

[not collecting (0 since last start), save path "data"]> size 1000000

refresh

Format:

> refresh <count>

Sets the number of collected tweets that triggers an automatic refresh of the prompt.

Example:

[not collecting (0 since last start), save path "data"]> refresh 10000

Implementing a custom command line tool

The TwiGetCLIBase class in twiget_cli.py module implements all the above fuctions except those related to saving to json file (i.e. save_to and size). It can be used to implement a command line tool that performs a different processing of the collected data, e.g., saving to a db.

Python class: TwiGet

TwiGet core functionalities are implemented in a python class, which can be directly used in python code.

fromtwigetimportTwiGetbearer='put here the bearer token'collector=TwiGet(bearer)
# Adding a filtering rule# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesquery='support vector machine'tag='ml'answer=collector.add_rule(query, tag)
# returns the parsed json answer from the server.print(answer)
# Listing the current filtering rules# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-stream-rulesanswer=collector.get_rules()
# returns the parsed json answer from the server.print(answer)
# Delete some rules by giving their ID# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesids= [48573094587309485,3029834285720978]
answer=collector.delete_rules(ids)
# returns the parsed json answer from the server.print(answer)
# Adding a callback# The data argument contains the content and information about the retrieved tweet# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-streamdefprint_tag(data):
print(data['matching_rules'][0]['tag'])
answer=collector.add_callback('print tag', print_tag)
# returns the parsed json answer from the server.print(answer)
# Getting callbackscallbacks=collector.get_callbacks()
# returns a list of pairs with the name of the callback and the callback method.print(callbacks)
# Delete a callbackcollector.delete_callback('print tag')
# Starting tweet collectioncollector.start_getting_stream()
# Checking status of collectionrunning=collector.is_getting_stream()
# returns a boolean. True if collection is active.print(running)
# Stopping tweet collectioncollector.stop_getting_stream()

License

Author: Andrea Esuli

BSD 3-Clause License, see license file

About

Get tweets from Twitter filtered stream API. Python and command line.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TwiGet

TwiGet is a python package for the management of the queries on filtered stream of the Twitter API, and the collection of tweets from it.

It can be used as a command line tool (twiget-cli) or as a python class (TwiGet).

Installation

> pip install twiget

The command installs the package and also makes the twiget-cli command available.

Command line tool: twiget-cli

TwiGet implements a command line interface that can be started with the command:

> twiget-cli

When launched without arguments the program searches for a .twiget.conf file in the HOME directory (the directory pointed by the $HOME or %userprofile% environment variable). The file must contain in the first line the bearer token that allows the program to access the Twitter API.

Alternatively, the name of the file from which to obtain the bearer token can be given as argument when starting the program:

> twiget-cli -b path_to_file/with_token.txt

NOTE: store the bearer token in a file with minimum access permissions. Never share it. Revoke any tokens that may have been made public.

Another optional argument is the path where to save collected tweets. By default, a data subdirectory is created in the current working directory.

> twiget-cli -s ./save_dir

prompt

When started, twiget-cli shown the available commands, and the queries currently registered for the given bearer token (queries are permanently stored on Twitter's servers).

TwiGet 0.1.1
Available commands (type help <command> for details):
create, delete, exit, help, list, refresh, save_to, size, start, stop
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=1405490304970434817	query="bts"	tag="bts"

The command prompt tells if twiget-cli is currently collecting tweets, the number of collected tweets, and the save path.

[not collecting (0 since last start), save path "data"]>

When collecting tweets, the prompt is automatically refreshed every time a given number of tweets is collected (see the refresh command).

Commands

create

Format:

> create <tag> <query>

Creates a filtering rule, associated to a given tag name.
Collected tweets are saved in json format in a file named <tag>.json, in the given save path. Tag name is the first argument of the command and cannot contain spaces.
Any word after the tag defines the query. Info on how to define rules.

Example:

[not collecting (0 since last start), save path "data"]>create usa jow biden
Tweets matching the query "jow biden" will be saved in the file data/usa.json
ID=1395720345987340524

list

Format:

> list

Lists the queries, their ID and their tag, currently registered for the filtered stream.

Example:

[not collecting (0 since last start), save path "data"]> list
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=13905490304970434817	query="bts"	tag="bts"
ID=1395720345987340524	query="joe biden"	tag="usa"

delete

Format:

> delete <ID>

Deletes a query, given its ID.

Example:

[not collecting (0 since last start), save path "data"]> delete 1385892384573355842

start

Format:

> start

Starts a background process that collects tweets from the filtered stream and puts them in json files, according to the tag they are associated to.

Collection continues until a stop or a exit command is entered. To let TwiGet collect data for longer periods of time, I suggest to use TwiGet within a virtual terminal session, using, e.g., screen or tmux.

Note: create and delete command can be issued also when collecting tweets. The collection process is updated immediately.

Example:

[not collecting (0 since last start), save path "data"]> start
[collecting (0 since last start), save path "data"]>

stop

Format:

> stop

Stop data collections.

Example:

[collecting (3000 since last start), save path "data"]> stop
[not collecting (3152 since last start), save path "data"]> 

save_to

Format:

> save_to <path>

Sets the path where json files are saved.

_Note: changing path while collecting tweets will immediately create new json file in the new path, leaving all tweets collected until that moment in the old path.

Example:

[not collecting (0 since last start), save path "data"]> save_to ../my_project
[not collecting (0 since last start), save path "../my_project"]> 

size

Format:

> size <size>

Sets the maximum size in bytes of json files. When a json file reaches this size, a new file with an incremented index (e.g., tag_0.json, tag_1.json, tag_2.json...) is created.

Example:

[not collecting (0 since last start), save path "data"]> size 1000000

refresh

Format:

> refresh <count>

Sets the number of collected tweets that triggers an automatic refresh of the prompt.

Example:

[not collecting (0 since last start), save path "data"]> refresh 10000

Implementing a custom command line tool

The TwiGetCLIBase class in twiget_cli.py module implements all the above fuctions except those related to saving to json file (i.e. save_to and size). It can be used to implement a command line tool that performs a different processing of the collected data, e.g., saving to a db.

Python class: TwiGet

TwiGet core functionalities are implemented in a python class, which can be directly used in python code.

fromtwigetimportTwiGetbearer='put here the bearer token'collector=TwiGet(bearer)
# Adding a filtering rule# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesquery='support vector machine'tag='ml'answer=collector.add_rule(query, tag)
# returns the parsed json answer from the server.print(answer)
# Listing the current filtering rules# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-stream-rulesanswer=collector.get_rules()
# returns the parsed json answer from the server.print(answer)
# Delete some rules by giving their ID# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesids= [48573094587309485,3029834285720978]
answer=collector.delete_rules(ids)
# returns the parsed json answer from the server.print(answer)
# Adding a callback# The data argument contains the content and information about the retrieved tweet# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-streamdefprint_tag(data):
print(data['matching_rules'][0]['tag'])
answer=collector.add_callback('print tag', print_tag)
# returns the parsed json answer from the server.print(answer)
# Getting callbackscallbacks=collector.get_callbacks()
# returns a list of pairs with the name of the callback and the callback method.print(callbacks)
# Delete a callbackcollector.delete_callback('print tag')
# Starting tweet collectioncollector.start_getting_stream()
# Checking status of collectionrunning=collector.is_getting_stream()
# returns a boolean. True if collection is active.print(running)
# Stopping tweet collectioncollector.stop_getting_stream()

License

Author: Andrea Esuli

BSD 3-Clause License, see license file

About

Get tweets from Twitter filtered stream API. Python and command line.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TwiGet

TwiGet is a python package for the management of the queries on filtered stream of the Twitter API, and the collection of tweets from it.

It can be used as a command line tool (twiget-cli) or as a python class (TwiGet).

Installation

> pip install twiget

The command installs the package and also makes the twiget-cli command available.

Command line tool: twiget-cli

TwiGet implements a command line interface that can be started with the command:

> twiget-cli

When launched without arguments the program searches for a .twiget.conf file in the HOME directory (the directory pointed by the $HOME or %userprofile% environment variable). The file must contain in the first line the bearer token that allows the program to access the Twitter API.

Alternatively, the name of the file from which to obtain the bearer token can be given as argument when starting the program:

> twiget-cli -b path_to_file/with_token.txt

NOTE: store the bearer token in a file with minimum access permissions. Never share it. Revoke any tokens that may have been made public.

Another optional argument is the path where to save collected tweets. By default, a data subdirectory is created in the current working directory.

> twiget-cli -s ./save_dir

prompt

When started, twiget-cli shown the available commands, and the queries currently registered for the given bearer token (queries are permanently stored on Twitter's servers).

TwiGet 0.1.1
Available commands (type help <command> for details):
create, delete, exit, help, list, refresh, save_to, size, start, stop
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=1405490304970434817	query="bts"	tag="bts"

The command prompt tells if twiget-cli is currently collecting tweets, the number of collected tweets, and the save path.

[not collecting (0 since last start), save path "data"]>

When collecting tweets, the prompt is automatically refreshed every time a given number of tweets is collected (see the refresh command).

Commands

create

Format:

> create <tag> <query>

Creates a filtering rule, associated to a given tag name.
Collected tweets are saved in json format in a file named <tag>.json, in the given save path. Tag name is the first argument of the command and cannot contain spaces.
Any word after the tag defines the query. Info on how to define rules.

Example:

[not collecting (0 since last start), save path "data"]>create usa jow biden
Tweets matching the query "jow biden" will be saved in the file data/usa.json
ID=1395720345987340524

list

Format:

> list

Lists the queries, their ID and their tag, currently registered for the filtered stream.

Example:

[not collecting (0 since last start), save path "data"]> list
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=13905490304970434817	query="bts"	tag="bts"
ID=1395720345987340524	query="joe biden"	tag="usa"

delete

Format:

> delete <ID>

Deletes a query, given its ID.

Example:

[not collecting (0 since last start), save path "data"]> delete 1385892384573355842

start

Format:

> start

Starts a background process that collects tweets from the filtered stream and puts them in json files, according to the tag they are associated to.

Collection continues until a stop or a exit command is entered. To let TwiGet collect data for longer periods of time, I suggest to use TwiGet within a virtual terminal session, using, e.g., screen or tmux.

Note: create and delete command can be issued also when collecting tweets. The collection process is updated immediately.

Example:

[not collecting (0 since last start), save path "data"]> start
[collecting (0 since last start), save path "data"]>

stop

Format:

> stop

Stop data collections.

Example:

[collecting (3000 since last start), save path "data"]> stop
[not collecting (3152 since last start), save path "data"]> 

save_to

Format:

> save_to <path>

Sets the path where json files are saved.

_Note: changing path while collecting tweets will immediately create new json file in the new path, leaving all tweets collected until that moment in the old path.

Example:

[not collecting (0 since last start), save path "data"]> save_to ../my_project
[not collecting (0 since last start), save path "../my_project"]> 

size

Format:

> size <size>

Sets the maximum size in bytes of json files. When a json file reaches this size, a new file with an incremented index (e.g., tag_0.json, tag_1.json, tag_2.json...) is created.

Example:

[not collecting (0 since last start), save path "data"]> size 1000000

refresh

Format:

> refresh <count>

Sets the number of collected tweets that triggers an automatic refresh of the prompt.

Example:

[not collecting (0 since last start), save path "data"]> refresh 10000

Implementing a custom command line tool

The TwiGetCLIBase class in twiget_cli.py module implements all the above fuctions except those related to saving to json file (i.e. save_to and size). It can be used to implement a command line tool that performs a different processing of the collected data, e.g., saving to a db.

Python class: TwiGet

TwiGet core functionalities are implemented in a python class, which can be directly used in python code.

fromtwigetimportTwiGetbearer='put here the bearer token'collector=TwiGet(bearer)
# Adding a filtering rule# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesquery='support vector machine'tag='ml'answer=collector.add_rule(query, tag)
# returns the parsed json answer from the server.print(answer)
# Listing the current filtering rules# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-stream-rulesanswer=collector.get_rules()
# returns the parsed json answer from the server.print(answer)
# Delete some rules by giving their ID# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesids= [48573094587309485,3029834285720978]
answer=collector.delete_rules(ids)
# returns the parsed json answer from the server.print(answer)
# Adding a callback# The data argument contains the content and information about the retrieved tweet# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-streamdefprint_tag(data):
print(data['matching_rules'][0]['tag'])
answer=collector.add_callback('print tag', print_tag)
# returns the parsed json answer from the server.print(answer)
# Getting callbackscallbacks=collector.get_callbacks()
# returns a list of pairs with the name of the callback and the callback method.print(callbacks)
# Delete a callbackcollector.delete_callback('print tag')
# Starting tweet collectioncollector.start_getting_stream()
# Checking status of collectionrunning=collector.is_getting_stream()
# returns a boolean. True if collection is active.print(running)
# Stopping tweet collectioncollector.stop_getting_stream()

License

Author: Andrea Esuli

BSD 3-Clause License, see license file

About

Get tweets from Twitter filtered stream API. Python and command line.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TwiGet

TwiGet is a python package for the management of the queries on filtered stream of the Twitter API, and the collection of tweets from it.

It can be used as a command line tool (twiget-cli) or as a python class (TwiGet).

Installation

> pip install twiget

The command installs the package and also makes the twiget-cli command available.

Command line tool: twiget-cli

TwiGet implements a command line interface that can be started with the command:

> twiget-cli

When launched without arguments the program searches for a .twiget.conf file in the HOME directory (the directory pointed by the $HOME or %userprofile% environment variable). The file must contain in the first line the bearer token that allows the program to access the Twitter API.

Alternatively, the name of the file from which to obtain the bearer token can be given as argument when starting the program:

> twiget-cli -b path_to_file/with_token.txt

NOTE: store the bearer token in a file with minimum access permissions. Never share it. Revoke any tokens that may have been made public.

Another optional argument is the path where to save collected tweets. By default, a data subdirectory is created in the current working directory.

> twiget-cli -s ./save_dir

prompt

When started, twiget-cli shown the available commands, and the queries currently registered for the given bearer token (queries are permanently stored on Twitter's servers).

TwiGet 0.1.1
Available commands (type help <command> for details):
create, delete, exit, help, list, refresh, save_to, size, start, stop
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=1405490304970434817	query="bts"	tag="bts"

The command prompt tells if twiget-cli is currently collecting tweets, the number of collected tweets, and the save path.

[not collecting (0 since last start), save path "data"]>

When collecting tweets, the prompt is automatically refreshed every time a given number of tweets is collected (see the refresh command).

Commands

create

Format:

> create <tag> <query>

Creates a filtering rule, associated to a given tag name.
Collected tweets are saved in json format in a file named <tag>.json, in the given save path. Tag name is the first argument of the command and cannot contain spaces.
Any word after the tag defines the query. Info on how to define rules.

Example:

[not collecting (0 since last start), save path "data"]>create usa jow biden
Tweets matching the query "jow biden" will be saved in the file data/usa.json
ID=1395720345987340524

list

Format:

> list

Lists the queries, their ID and their tag, currently registered for the filtered stream.

Example:

[not collecting (0 since last start), save path "data"]> list
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=13905490304970434817	query="bts"	tag="bts"
ID=1395720345987340524	query="joe biden"	tag="usa"

delete

Format:

> delete <ID>

Deletes a query, given its ID.

Example:

[not collecting (0 since last start), save path "data"]> delete 1385892384573355842

start

Format:

> start

Starts a background process that collects tweets from the filtered stream and puts them in json files, according to the tag they are associated to.

Collection continues until a stop or a exit command is entered. To let TwiGet collect data for longer periods of time, I suggest to use TwiGet within a virtual terminal session, using, e.g., screen or tmux.

Note: create and delete command can be issued also when collecting tweets. The collection process is updated immediately.

Example:

[not collecting (0 since last start), save path "data"]> start
[collecting (0 since last start), save path "data"]>

stop

Format:

> stop

Stop data collections.

Example:

[collecting (3000 since last start), save path "data"]> stop
[not collecting (3152 since last start), save path "data"]> 

save_to

Format:

> save_to <path>

Sets the path where json files are saved.

_Note: changing path while collecting tweets will immediately create new json file in the new path, leaving all tweets collected until that moment in the old path.

Example:

[not collecting (0 since last start), save path "data"]> save_to ../my_project
[not collecting (0 since last start), save path "../my_project"]> 

size

Format:

> size <size>

Sets the maximum size in bytes of json files. When a json file reaches this size, a new file with an incremented index (e.g., tag_0.json, tag_1.json, tag_2.json...) is created.

Example:

[not collecting (0 since last start), save path "data"]> size 1000000

refresh

Format:

> refresh <count>

Sets the number of collected tweets that triggers an automatic refresh of the prompt.

Example:

[not collecting (0 since last start), save path "data"]> refresh 10000

Implementing a custom command line tool

The TwiGetCLIBase class in twiget_cli.py module implements all the above fuctions except those related to saving to json file (i.e. save_to and size). It can be used to implement a command line tool that performs a different processing of the collected data, e.g., saving to a db.

Python class: TwiGet

TwiGet core functionalities are implemented in a python class, which can be directly used in python code.

fromtwigetimportTwiGetbearer='put here the bearer token'collector=TwiGet(bearer)
# Adding a filtering rule# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesquery='support vector machine'tag='ml'answer=collector.add_rule(query, tag)
# returns the parsed json answer from the server.print(answer)
# Listing the current filtering rules# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-stream-rulesanswer=collector.get_rules()
# returns the parsed json answer from the server.print(answer)
# Delete some rules by giving their ID# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesids= [48573094587309485,3029834285720978]
answer=collector.delete_rules(ids)
# returns the parsed json answer from the server.print(answer)
# Adding a callback# The data argument contains the content and information about the retrieved tweet# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-streamdefprint_tag(data):
print(data['matching_rules'][0]['tag'])
answer=collector.add_callback('print tag', print_tag)
# returns the parsed json answer from the server.print(answer)
# Getting callbackscallbacks=collector.get_callbacks()
# returns a list of pairs with the name of the callback and the callback method.print(callbacks)
# Delete a callbackcollector.delete_callback('print tag')
# Starting tweet collectioncollector.start_getting_stream()
# Checking status of collectionrunning=collector.is_getting_stream()
# returns a boolean. True if collection is active.print(running)
# Stopping tweet collectioncollector.stop_getting_stream()

License

Author: Andrea Esuli

BSD 3-Clause License, see license file

About

Get tweets from Twitter filtered stream API. Python and command line.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TwiGet

TwiGet is a python package for the management of the queries on filtered stream of the Twitter API, and the collection of tweets from it.

It can be used as a command line tool (twiget-cli) or as a python class (TwiGet).

Installation

> pip install twiget

The command installs the package and also makes the twiget-cli command available.

Command line tool: twiget-cli

TwiGet implements a command line interface that can be started with the command:

> twiget-cli

When launched without arguments the program searches for a .twiget.conf file in the HOME directory (the directory pointed by the $HOME or %userprofile% environment variable). The file must contain in the first line the bearer token that allows the program to access the Twitter API.

Alternatively, the name of the file from which to obtain the bearer token can be given as argument when starting the program:

> twiget-cli -b path_to_file/with_token.txt

NOTE: store the bearer token in a file with minimum access permissions. Never share it. Revoke any tokens that may have been made public.

Another optional argument is the path where to save collected tweets. By default, a data subdirectory is created in the current working directory.

> twiget-cli -s ./save_dir

prompt

When started, twiget-cli shown the available commands, and the queries currently registered for the given bearer token (queries are permanently stored on Twitter's servers).

TwiGet 0.1.1
Available commands (type help <command> for details):
create, delete, exit, help, list, refresh, save_to, size, start, stop
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=1405490304970434817	query="bts"	tag="bts"

The command prompt tells if twiget-cli is currently collecting tweets, the number of collected tweets, and the save path.

[not collecting (0 since last start), save path "data"]>

When collecting tweets, the prompt is automatically refreshed every time a given number of tweets is collected (see the refresh command).

Commands

create

Format:

> create <tag> <query>

Creates a filtering rule, associated to a given tag name.
Collected tweets are saved in json format in a file named <tag>.json, in the given save path. Tag name is the first argument of the command and cannot contain spaces.
Any word after the tag defines the query. Info on how to define rules.

Example:

[not collecting (0 since last start), save path "data"]>create usa jow biden
Tweets matching the query "jow biden" will be saved in the file data/usa.json
ID=1395720345987340524

list

Format:

> list

Lists the queries, their ID and their tag, currently registered for the filtered stream.

Example:

[not collecting (0 since last start), save path "data"]> list
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=13905490304970434817	query="bts"	tag="bts"
ID=1395720345987340524	query="joe biden"	tag="usa"

delete

Format:

> delete <ID>

Deletes a query, given its ID.

Example:

[not collecting (0 since last start), save path "data"]> delete 1385892384573355842

start

Format:

> start

Starts a background process that collects tweets from the filtered stream and puts them in json files, according to the tag they are associated to.

Collection continues until a stop or a exit command is entered. To let TwiGet collect data for longer periods of time, I suggest to use TwiGet within a virtual terminal session, using, e.g., screen or tmux.

Note: create and delete command can be issued also when collecting tweets. The collection process is updated immediately.

Example:

[not collecting (0 since last start), save path "data"]> start
[collecting (0 since last start), save path "data"]>

stop

Format:

> stop

Stop data collections.

Example:

[collecting (3000 since last start), save path "data"]> stop
[not collecting (3152 since last start), save path "data"]> 

save_to

Format:

> save_to <path>

Sets the path where json files are saved.

_Note: changing path while collecting tweets will immediately create new json file in the new path, leaving all tweets collected until that moment in the old path.

Example:

[not collecting (0 since last start), save path "data"]> save_to ../my_project
[not collecting (0 since last start), save path "../my_project"]> 

size

Format:

> size <size>

Sets the maximum size in bytes of json files. When a json file reaches this size, a new file with an incremented index (e.g., tag_0.json, tag_1.json, tag_2.json...) is created.

Example:

[not collecting (0 since last start), save path "data"]> size 1000000

refresh

Format:

> refresh <count>

Sets the number of collected tweets that triggers an automatic refresh of the prompt.

Example:

[not collecting (0 since last start), save path "data"]> refresh 10000

Implementing a custom command line tool

The TwiGetCLIBase class in twiget_cli.py module implements all the above fuctions except those related to saving to json file (i.e. save_to and size). It can be used to implement a command line tool that performs a different processing of the collected data, e.g., saving to a db.

Python class: TwiGet

TwiGet core functionalities are implemented in a python class, which can be directly used in python code.

fromtwigetimportTwiGetbearer='put here the bearer token'collector=TwiGet(bearer)
# Adding a filtering rule# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesquery='support vector machine'tag='ml'answer=collector.add_rule(query, tag)
# returns the parsed json answer from the server.print(answer)
# Listing the current filtering rules# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-stream-rulesanswer=collector.get_rules()
# returns the parsed json answer from the server.print(answer)
# Delete some rules by giving their ID# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesids= [48573094587309485,3029834285720978]
answer=collector.delete_rules(ids)
# returns the parsed json answer from the server.print(answer)
# Adding a callback# The data argument contains the content and information about the retrieved tweet# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-streamdefprint_tag(data):
print(data['matching_rules'][0]['tag'])
answer=collector.add_callback('print tag', print_tag)
# returns the parsed json answer from the server.print(answer)
# Getting callbackscallbacks=collector.get_callbacks()
# returns a list of pairs with the name of the callback and the callback method.print(callbacks)
# Delete a callbackcollector.delete_callback('print tag')
# Starting tweet collectioncollector.start_getting_stream()
# Checking status of collectionrunning=collector.is_getting_stream()
# returns a boolean. True if collection is active.print(running)
# Stopping tweet collectioncollector.stop_getting_stream()

License

Author: Andrea Esuli

BSD 3-Clause License, see license file

About

Get tweets from Twitter filtered stream API. Python and command line.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TwiGet

TwiGet is a python package for the management of the queries on filtered stream of the Twitter API, and the collection of tweets from it.

It can be used as a command line tool (twiget-cli) or as a python class (TwiGet).

Installation

> pip install twiget

The command installs the package and also makes the twiget-cli command available.

Command line tool: twiget-cli

TwiGet implements a command line interface that can be started with the command:

> twiget-cli

When launched without arguments the program searches for a .twiget.conf file in the HOME directory (the directory pointed by the $HOME or %userprofile% environment variable). The file must contain in the first line the bearer token that allows the program to access the Twitter API.

Alternatively, the name of the file from which to obtain the bearer token can be given as argument when starting the program:

> twiget-cli -b path_to_file/with_token.txt

NOTE: store the bearer token in a file with minimum access permissions. Never share it. Revoke any tokens that may have been made public.

Another optional argument is the path where to save collected tweets. By default, a data subdirectory is created in the current working directory.

> twiget-cli -s ./save_dir

prompt

When started, twiget-cli shown the available commands, and the queries currently registered for the given bearer token (queries are permanently stored on Twitter's servers).

TwiGet 0.1.1
Available commands (type help <command> for details):
create, delete, exit, help, list, refresh, save_to, size, start, stop
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=1405490304970434817	query="bts"	tag="bts"

The command prompt tells if twiget-cli is currently collecting tweets, the number of collected tweets, and the save path.

[not collecting (0 since last start), save path "data"]>

When collecting tweets, the prompt is automatically refreshed every time a given number of tweets is collected (see the refresh command).

Commands

create

Format:

> create <tag> <query>

Creates a filtering rule, associated to a given tag name.
Collected tweets are saved in json format in a file named <tag>.json, in the given save path. Tag name is the first argument of the command and cannot contain spaces.
Any word after the tag defines the query. Info on how to define rules.

Example:

[not collecting (0 since last start), save path "data"]>create usa jow biden
Tweets matching the query "jow biden" will be saved in the file data/usa.json
ID=1395720345987340524

list

Format:

> list

Lists the queries, their ID and their tag, currently registered for the filtered stream.

Example:

[not collecting (0 since last start), save path "data"]> list
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=13905490304970434817	query="bts"	tag="bts"
ID=1395720345987340524	query="joe biden"	tag="usa"

delete

Format:

> delete <ID>

Deletes a query, given its ID.

Example:

[not collecting (0 since last start), save path "data"]> delete 1385892384573355842

start

Format:

> start

Starts a background process that collects tweets from the filtered stream and puts them in json files, according to the tag they are associated to.

Collection continues until a stop or a exit command is entered. To let TwiGet collect data for longer periods of time, I suggest to use TwiGet within a virtual terminal session, using, e.g., screen or tmux.

Note: create and delete command can be issued also when collecting tweets. The collection process is updated immediately.

Example:

[not collecting (0 since last start), save path "data"]> start
[collecting (0 since last start), save path "data"]>

stop

Format:

> stop

Stop data collections.

Example:

[collecting (3000 since last start), save path "data"]> stop
[not collecting (3152 since last start), save path "data"]> 

save_to

Format:

> save_to <path>

Sets the path where json files are saved.

_Note: changing path while collecting tweets will immediately create new json file in the new path, leaving all tweets collected until that moment in the old path.

Example:

[not collecting (0 since last start), save path "data"]> save_to ../my_project
[not collecting (0 since last start), save path "../my_project"]> 

size

Format:

> size <size>

Sets the maximum size in bytes of json files. When a json file reaches this size, a new file with an incremented index (e.g., tag_0.json, tag_1.json, tag_2.json...) is created.

Example:

[not collecting (0 since last start), save path "data"]> size 1000000

refresh

Format:

> refresh <count>

Sets the number of collected tweets that triggers an automatic refresh of the prompt.

Example:

[not collecting (0 since last start), save path "data"]> refresh 10000

Implementing a custom command line tool

The TwiGetCLIBase class in twiget_cli.py module implements all the above fuctions except those related to saving to json file (i.e. save_to and size). It can be used to implement a command line tool that performs a different processing of the collected data, e.g., saving to a db.

Python class: TwiGet

TwiGet core functionalities are implemented in a python class, which can be directly used in python code.

fromtwigetimportTwiGetbearer='put here the bearer token'collector=TwiGet(bearer)
# Adding a filtering rule# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesquery='support vector machine'tag='ml'answer=collector.add_rule(query, tag)
# returns the parsed json answer from the server.print(answer)
# Listing the current filtering rules# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-stream-rulesanswer=collector.get_rules()
# returns the parsed json answer from the server.print(answer)
# Delete some rules by giving their ID# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesids= [48573094587309485,3029834285720978]
answer=collector.delete_rules(ids)
# returns the parsed json answer from the server.print(answer)
# Adding a callback# The data argument contains the content and information about the retrieved tweet# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-streamdefprint_tag(data):
print(data['matching_rules'][0]['tag'])
answer=collector.add_callback('print tag', print_tag)
# returns the parsed json answer from the server.print(answer)
# Getting callbackscallbacks=collector.get_callbacks()
# returns a list of pairs with the name of the callback and the callback method.print(callbacks)
# Delete a callbackcollector.delete_callback('print tag')
# Starting tweet collectioncollector.start_getting_stream()
# Checking status of collectionrunning=collector.is_getting_stream()
# returns a boolean. True if collection is active.print(running)
# Stopping tweet collectioncollector.stop_getting_stream()

License

Author: Andrea Esuli

BSD 3-Clause License, see license file

About

Get tweets from Twitter filtered stream API. Python and command line.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TwiGet

TwiGet is a python package for the management of the queries on filtered stream of the Twitter API, and the collection of tweets from it.

It can be used as a command line tool (twiget-cli) or as a python class (TwiGet).

Installation

> pip install twiget

The command installs the package and also makes the twiget-cli command available.

Command line tool: twiget-cli

TwiGet implements a command line interface that can be started with the command:

> twiget-cli

When launched without arguments the program searches for a .twiget.conf file in the HOME directory (the directory pointed by the $HOME or %userprofile% environment variable). The file must contain in the first line the bearer token that allows the program to access the Twitter API.

Alternatively, the name of the file from which to obtain the bearer token can be given as argument when starting the program:

> twiget-cli -b path_to_file/with_token.txt

NOTE: store the bearer token in a file with minimum access permissions. Never share it. Revoke any tokens that may have been made public.

Another optional argument is the path where to save collected tweets. By default, a data subdirectory is created in the current working directory.

> twiget-cli -s ./save_dir

prompt

When started, twiget-cli shown the available commands, and the queries currently registered for the given bearer token (queries are permanently stored on Twitter's servers).

TwiGet 0.1.1
Available commands (type help <command> for details):
create, delete, exit, help, list, refresh, save_to, size, start, stop
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=1405490304970434817	query="bts"	tag="bts"

The command prompt tells if twiget-cli is currently collecting tweets, the number of collected tweets, and the save path.

[not collecting (0 since last start), save path "data"]>

When collecting tweets, the prompt is automatically refreshed every time a given number of tweets is collected (see the refresh command).

Commands

create

Format:

> create <tag> <query>

Creates a filtering rule, associated to a given tag name.
Collected tweets are saved in json format in a file named <tag>.json, in the given save path. Tag name is the first argument of the command and cannot contain spaces.
Any word after the tag defines the query. Info on how to define rules.

Example:

[not collecting (0 since last start), save path "data"]>create usa jow biden
Tweets matching the query "jow biden" will be saved in the file data/usa.json
ID=1395720345987340524

list

Format:

> list

Lists the queries, their ID and their tag, currently registered for the filtered stream.

Example:

[not collecting (0 since last start), save path "data"]> list
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=13905490304970434817	query="bts"	tag="bts"
ID=1395720345987340524	query="joe biden"	tag="usa"

delete

Format:

> delete <ID>

Deletes a query, given its ID.

Example:

[not collecting (0 since last start), save path "data"]> delete 1385892384573355842

start

Format:

> start

Starts a background process that collects tweets from the filtered stream and puts them in json files, according to the tag they are associated to.

Collection continues until a stop or a exit command is entered. To let TwiGet collect data for longer periods of time, I suggest to use TwiGet within a virtual terminal session, using, e.g., screen or tmux.

Note: create and delete command can be issued also when collecting tweets. The collection process is updated immediately.

Example:

[not collecting (0 since last start), save path "data"]> start
[collecting (0 since last start), save path "data"]>

stop

Format:

> stop

Stop data collections.

Example:

[collecting (3000 since last start), save path "data"]> stop
[not collecting (3152 since last start), save path "data"]> 

save_to

Format:

> save_to <path>

Sets the path where json files are saved.

_Note: changing path while collecting tweets will immediately create new json file in the new path, leaving all tweets collected until that moment in the old path.

Example:

[not collecting (0 since last start), save path "data"]> save_to ../my_project
[not collecting (0 since last start), save path "../my_project"]> 

size

Format:

> size <size>

Sets the maximum size in bytes of json files. When a json file reaches this size, a new file with an incremented index (e.g., tag_0.json, tag_1.json, tag_2.json...) is created.

Example:

[not collecting (0 since last start), save path "data"]> size 1000000

refresh

Format:

> refresh <count>

Sets the number of collected tweets that triggers an automatic refresh of the prompt.

Example:

[not collecting (0 since last start), save path "data"]> refresh 10000

Implementing a custom command line tool

The TwiGetCLIBase class in twiget_cli.py module implements all the above fuctions except those related to saving to json file (i.e. save_to and size). It can be used to implement a command line tool that performs a different processing of the collected data, e.g., saving to a db.

Python class: TwiGet

TwiGet core functionalities are implemented in a python class, which can be directly used in python code.

fromtwigetimportTwiGetbearer='put here the bearer token'collector=TwiGet(bearer)
# Adding a filtering rule# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesquery='support vector machine'tag='ml'answer=collector.add_rule(query, tag)
# returns the parsed json answer from the server.print(answer)
# Listing the current filtering rules# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-stream-rulesanswer=collector.get_rules()
# returns the parsed json answer from the server.print(answer)
# Delete some rules by giving their ID# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesids= [48573094587309485,3029834285720978]
answer=collector.delete_rules(ids)
# returns the parsed json answer from the server.print(answer)
# Adding a callback# The data argument contains the content and information about the retrieved tweet# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-streamdefprint_tag(data):
print(data['matching_rules'][0]['tag'])
answer=collector.add_callback('print tag', print_tag)
# returns the parsed json answer from the server.print(answer)
# Getting callbackscallbacks=collector.get_callbacks()
# returns a list of pairs with the name of the callback and the callback method.print(callbacks)
# Delete a callbackcollector.delete_callback('print tag')
# Starting tweet collectioncollector.start_getting_stream()
# Checking status of collectionrunning=collector.is_getting_stream()
# returns a boolean. True if collection is active.print(running)
# Stopping tweet collectioncollector.stop_getting_stream()

License

Author: Andrea Esuli

BSD 3-Clause License, see license file

About

Get tweets from Twitter filtered stream API. Python and command line.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

TwiGet

TwiGet is a python package for the management of the queries on filtered stream of the Twitter API, and the collection of tweets from it.

It can be used as a command line tool (twiget-cli) or as a python class (TwiGet).

Installation

> pip install twiget

The command installs the package and also makes the twiget-cli command available.

Command line tool: twiget-cli

TwiGet implements a command line interface that can be started with the command:

> twiget-cli

When launched without arguments the program searches for a .twiget.conf file in the HOME directory (the directory pointed by the $HOME or %userprofile% environment variable). The file must contain in the first line the bearer token that allows the program to access the Twitter API.

Alternatively, the name of the file from which to obtain the bearer token can be given as argument when starting the program:

> twiget-cli -b path_to_file/with_token.txt

NOTE: store the bearer token in a file with minimum access permissions. Never share it. Revoke any tokens that may have been made public.

Another optional argument is the path where to save collected tweets. By default, a data subdirectory is created in the current working directory.

> twiget-cli -s ./save_dir

prompt

When started, twiget-cli shown the available commands, and the queries currently registered for the given bearer token (queries are permanently stored on Twitter's servers).

TwiGet 0.1.1
Available commands (type help <command> for details):
create, delete, exit, help, list, refresh, save_to, size, start, stop
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=1405490304970434817	query="bts"	tag="bts"

The command prompt tells if twiget-cli is currently collecting tweets, the number of collected tweets, and the save path.

[not collecting (0 since last start), save path "data"]>

When collecting tweets, the prompt is automatically refreshed every time a given number of tweets is collected (see the refresh command).

Commands

create

Format:

> create <tag> <query>

Creates a filtering rule, associated to a given tag name.
Collected tweets are saved in json format in a file named <tag>.json, in the given save path. Tag name is the first argument of the command and cannot contain spaces.
Any word after the tag defines the query. Info on how to define rules.

Example:

[not collecting (0 since last start), save path "data"]>create usa jow biden
Tweets matching the query "jow biden" will be saved in the file data/usa.json
ID=1395720345987340524

list

Format:

> list

Lists the queries, their ID and their tag, currently registered for the filtered stream.

Example:

[not collecting (0 since last start), save path "data"]> list
Registered queries:
ID=1385892384573355842	query="#usa"	tag="usa"
ID=13905490304970434817	query="bts"	tag="bts"
ID=1395720345987340524	query="joe biden"	tag="usa"

delete

Format:

> delete <ID>

Deletes a query, given its ID.

Example:

[not collecting (0 since last start), save path "data"]> delete 1385892384573355842

start

Format:

> start

Starts a background process that collects tweets from the filtered stream and puts them in json files, according to the tag they are associated to.

Collection continues until a stop or a exit command is entered. To let TwiGet collect data for longer periods of time, I suggest to use TwiGet within a virtual terminal session, using, e.g., screen or tmux.

Note: create and delete command can be issued also when collecting tweets. The collection process is updated immediately.

Example:

[not collecting (0 since last start), save path "data"]> start
[collecting (0 since last start), save path "data"]>

stop

Format:

> stop

Stop data collections.

Example:

[collecting (3000 since last start), save path "data"]> stop
[not collecting (3152 since last start), save path "data"]> 

save_to

Format:

> save_to <path>

Sets the path where json files are saved.

_Note: changing path while collecting tweets will immediately create new json file in the new path, leaving all tweets collected until that moment in the old path.

Example:

[not collecting (0 since last start), save path "data"]> save_to ../my_project
[not collecting (0 since last start), save path "../my_project"]> 

size

Format:

> size <size>

Sets the maximum size in bytes of json files. When a json file reaches this size, a new file with an incremented index (e.g., tag_0.json, tag_1.json, tag_2.json...) is created.

Example:

[not collecting (0 since last start), save path "data"]> size 1000000

refresh

Format:

> refresh <count>

Sets the number of collected tweets that triggers an automatic refresh of the prompt.

Example:

[not collecting (0 since last start), save path "data"]> refresh 10000

Implementing a custom command line tool

The TwiGetCLIBase class in twiget_cli.py module implements all the above fuctions except those related to saving to json file (i.e. save_to and size). It can be used to implement a command line tool that performs a different processing of the collected data, e.g., saving to a db.

Python class: TwiGet

TwiGet core functionalities are implemented in a python class, which can be directly used in python code.

fromtwigetimportTwiGetbearer='put here the bearer token'collector=TwiGet(bearer)
# Adding a filtering rule# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesquery='support vector machine'tag='ml'answer=collector.add_rule(query, tag)
# returns the parsed json answer from the server.print(answer)
# Listing the current filtering rules# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-stream-rulesanswer=collector.get_rules()
# returns the parsed json answer from the server.print(answer)
# Delete some rules by giving their ID# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/post-tweets-search-stream-rulesids= [48573094587309485,3029834285720978]
answer=collector.delete_rules(ids)
# returns the parsed json answer from the server.print(answer)
# Adding a callback# The data argument contains the content and information about the retrieved tweet# https://developer.twitter.com/en/docs/twitter-api/tweets/filtered-stream/api-reference/get-tweets-search-streamdefprint_tag(data):
print(data['matching_rules'][0]['tag'])
answer=collector.add_callback('print tag', print_tag)
# returns the parsed json answer from the server.print(answer)
# Getting callbackscallbacks=collector.get_callbacks()
# returns a list of pairs with the name of the callback and the callback method.print(callbacks)
# Delete a callbackcollector.delete_callback('print tag')
# Starting tweet collectioncollector.start_getting_stream()
# Checking status of collectionrunning=collector.is_getting_stream()
# returns a boolean. True if collection is active.print(running)
# Stopping tweet collectioncollector.stop_getting_stream()

License

Author: Andrea Esuli

BSD 3-Clause License, see license file

About

Get tweets from Twitter filtered stream API. Python and command line.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages