The Ollama Python library provides the easiest way to integrate Python 3.8+ projects with Ollama.
pip install ollamaimportollamaresponse=ollama.chat(model='llama2', messages=[
{
'role': 'user',
'content': 'Why is the sky blue?',
},
])
print(response['message']['content'])Response streaming can be enabled by setting stream=True, modifying function calls to return a Python generator where each part is an object in the stream.
importollamastream=ollama.chat(
model='llama2',
messages=[{'role': 'user', 'content': 'Why is the sky blue?'}],
stream=True,
)
forchunkinstream:
print(chunk['message']['content'], end='', flush=True)The Ollama Python library's API is designed around the Ollama REST API
ollama.chat(model='llama2', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}])ollama.generate(model='llama2', prompt='Why is the sky blue?')ollama.list()ollama.show('llama2')modelfile='''FROM llama2SYSTEM You are mario from super mario bros.'''ollama.create(model='example', modelfile=modelfile)ollama.copy('llama2', 'user/llama2')ollama.delete('llama2')ollama.pull('llama2')ollama.push('user/llama2')ollama.embeddings(model='llama2', prompt='They sky is blue because of rayleigh scattering')A custom client can be created with the following fields:
host: The Ollama host to connect totimeout: The timeout for requests
fromollamaimportClientclient=Client(host='http://localhost:11434')
response=client.chat(model='llama2', messages=[
{
'role': 'user',
'content': 'Why is the sky blue?',
},
])importasynciofromollamaimportAsyncClientasyncdefchat():
message= {'role': 'user', 'content': 'Why is the sky blue?'}
response=awaitAsyncClient().chat(model='llama2', messages=[message])
asyncio.run(chat())Setting stream=True modifies functions to return a Python asynchronous generator:
importasynciofromollamaimportAsyncClientasyncdefchat():
message= {'role': 'user', 'content': 'Why is the sky blue?'}
asyncforpartinawaitAsyncClient().chat(model='llama2', messages=[message], stream=True):
print(part['message']['content'], end='', flush=True)
asyncio.run(chat())Errors are raised if requests return an error status or if an error is detected while streaming.
model='does-not-yet-exist'try:
ollama.chat(model)
exceptollama.ResponseErrorase:
print('Error:', e.error)
ife.status_code==404:
ollama.pull(model)