Repository files navigation

Chainference

Decentralized inference on the Solana chain.

As a server, you can monetize your hardware by running paid AI inference for users.

As a user, you get access to open AI models without restrictions and for cheaper than on centralized services.

Cheaper because unrestricted competition drives prices down to healthy profit margins.

Getting started

See each subfolder's readme for instructions on the different parts of the project.

How it works

  1. Servers publish availability on-chain
  2. Clients publish inference requests on-chain, staking maximum desired cost
  3. A server locks the inference request on-chain
  4. Client sends prompt to server off-chain
  5. Server streams response to client off-chain
  6. Server claims payment on-chain, with remainder returned to client

1. Servers publish availability on-chain

Servers publish availability by sending a transaction to a smart contract, including a list of models available for inference (hugging face IDs) and price per million output tokens for each.

2. Clients publish inference requests on-chain

Clients publish inference requests on chain also by sending a transaction to a smart contract.

They don't send requests directly to servers because the maximum cost must be staked on-chain to ensure payment.

An inference request includes:

  • Filtering criteria for server
  • Desired model
  • SOL stake of maximum cost
  • Creation timestamp

3. A server locks the inference request on-chain

Servers lock inference requests, once again, by sending a transaction to a smart contract.

This transaction includes the peer to peer address the client should send the prompt to.

4. Client sends prompt to server off-chain

The client then sends the prompt to the provided peer to peer address.

This should be encrypted - there may be encryption built into the peer to peer library or protocol we end up using, otherwise we may have client and server sharing public keys during the request submitting and request locking process.

5. Server streams response to client off-chain

After receiving the prompt, the server streams the response to the peer to peer address previously provided by the client.

This response, similarly to the prompt, should also be encrypted.

6. Server claims payment on-chain

When a server finishes responding to an inference request, it can charge the corresponding cost from the amount previously staked by the client, again by submitting a transaction to a smart contract.

After locking an inference request, servers have a 1 hour limit to claim payment. If they don't, the request is cancelled and the stake is fully returned to the user.

This should be a rare case, as there is no incentive for servers to not claim payment on a request, malicious or not.

This limit does handle the case where a benevolent server fails to respond to a request. The server may still get flagged by the client regardless, and likely before this time limit is reached, affecting its reputation.

Reputation system

Malicious behavior is disincentivized through a reputation system.

Monetization

We can introduce a fee (maybe around 10%) on the cost of each inference through the smart contract at some point in the future.

Question: should this be charged from the client's stake in addition to the server's charge, or deducted from the server's charge? Maybe the latter, as servers are the ones profiting so they should be the ones deducting.

Storage

TODO

Could be IPFS/filecoin or arweave

Website

Chainference will have a centralized website with:

  • List of active servers and their details
  • List of models and their details such as pricing and server availability

Solana facts

  • Fees are cheap, with $0.01 being enough for around 8k transactions at the base fee with a SOL value of $250. Prioritization fees are almost negligible at an average of ~1 lamport (1e-9 SOL).
  • Storage is expensive, with a cost of ~7 SOL per MB ($1750 at $250 SOL price). This value is however fully recoverable on deletion.

Thoughts on privacy

Anonymity

Prompts and responses are only visible to each client and server engaging in a transaction.

However, anonymity could be enhanced by using disposable accounts to submit prompts with.

To avoid tracing back to the real wallet, a centralized place could be creating these disposable wallets for everyone, perhaps the chainference smart contract itself.

However, servers can still associate prompts to the address they are sending the response to.

So the possibility of users creating temporary addresses to receive responses at should be investigated - some sort of proxying.

Encryption

Future improvements in encrypted inference may allow us to have servers running inference on prompts without being able to see its plain text contents, and perhaps also for the response.

About

LLM inference on the solana chain.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Chainference

Decentralized inference on the Solana chain.

As a server, you can monetize your hardware by running paid AI inference for users.

As a user, you get access to open AI models without restrictions and for cheaper than on centralized services.

Cheaper because unrestricted competition drives prices down to healthy profit margins.

Getting started

See each subfolder's readme for instructions on the different parts of the project.

How it works

  1. Servers publish availability on-chain
  2. Clients publish inference requests on-chain, staking maximum desired cost
  3. A server locks the inference request on-chain
  4. Client sends prompt to server off-chain
  5. Server streams response to client off-chain
  6. Server claims payment on-chain, with remainder returned to client

1. Servers publish availability on-chain

Servers publish availability by sending a transaction to a smart contract, including a list of models available for inference (hugging face IDs) and price per million output tokens for each.

2. Clients publish inference requests on-chain

Clients publish inference requests on chain also by sending a transaction to a smart contract.

They don't send requests directly to servers because the maximum cost must be staked on-chain to ensure payment.

An inference request includes:

  • Filtering criteria for server
  • Desired model
  • SOL stake of maximum cost
  • Creation timestamp

3. A server locks the inference request on-chain

Servers lock inference requests, once again, by sending a transaction to a smart contract.

This transaction includes the peer to peer address the client should send the prompt to.

4. Client sends prompt to server off-chain

The client then sends the prompt to the provided peer to peer address.

This should be encrypted - there may be encryption built into the peer to peer library or protocol we end up using, otherwise we may have client and server sharing public keys during the request submitting and request locking process.

5. Server streams response to client off-chain

After receiving the prompt, the server streams the response to the peer to peer address previously provided by the client.

This response, similarly to the prompt, should also be encrypted.

6. Server claims payment on-chain

When a server finishes responding to an inference request, it can charge the corresponding cost from the amount previously staked by the client, again by submitting a transaction to a smart contract.

After locking an inference request, servers have a 1 hour limit to claim payment. If they don't, the request is cancelled and the stake is fully returned to the user.

This should be a rare case, as there is no incentive for servers to not claim payment on a request, malicious or not.

This limit does handle the case where a benevolent server fails to respond to a request. The server may still get flagged by the client regardless, and likely before this time limit is reached, affecting its reputation.

Reputation system

Malicious behavior is disincentivized through a reputation system.

Monetization

We can introduce a fee (maybe around 10%) on the cost of each inference through the smart contract at some point in the future.

Question: should this be charged from the client's stake in addition to the server's charge, or deducted from the server's charge? Maybe the latter, as servers are the ones profiting so they should be the ones deducting.

Storage

TODO

Could be IPFS/filecoin or arweave

Website

Chainference will have a centralized website with:

  • List of active servers and their details
  • List of models and their details such as pricing and server availability

Solana facts

  • Fees are cheap, with $0.01 being enough for around 8k transactions at the base fee with a SOL value of $250. Prioritization fees are almost negligible at an average of ~1 lamport (1e-9 SOL).
  • Storage is expensive, with a cost of ~7 SOL per MB ($1750 at $250 SOL price). This value is however fully recoverable on deletion.

Thoughts on privacy

Anonymity

Prompts and responses are only visible to each client and server engaging in a transaction.

However, anonymity could be enhanced by using disposable accounts to submit prompts with.

To avoid tracing back to the real wallet, a centralized place could be creating these disposable wallets for everyone, perhaps the chainference smart contract itself.

However, servers can still associate prompts to the address they are sending the response to.

So the possibility of users creating temporary addresses to receive responses at should be investigated - some sort of proxying.

Encryption

Future improvements in encrypted inference may allow us to have servers running inference on prompts without being able to see its plain text contents, and perhaps also for the response.

About

LLM inference on the solana chain.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Chainference

Decentralized inference on the Solana chain.

As a server, you can monetize your hardware by running paid AI inference for users.

As a user, you get access to open AI models without restrictions and for cheaper than on centralized services.

Cheaper because unrestricted competition drives prices down to healthy profit margins.

Getting started

See each subfolder's readme for instructions on the different parts of the project.

How it works

  1. Servers publish availability on-chain
  2. Clients publish inference requests on-chain, staking maximum desired cost
  3. A server locks the inference request on-chain
  4. Client sends prompt to server off-chain
  5. Server streams response to client off-chain
  6. Server claims payment on-chain, with remainder returned to client

1. Servers publish availability on-chain

Servers publish availability by sending a transaction to a smart contract, including a list of models available for inference (hugging face IDs) and price per million output tokens for each.

2. Clients publish inference requests on-chain

Clients publish inference requests on chain also by sending a transaction to a smart contract.

They don't send requests directly to servers because the maximum cost must be staked on-chain to ensure payment.

An inference request includes:

  • Filtering criteria for server
  • Desired model
  • SOL stake of maximum cost
  • Creation timestamp

3. A server locks the inference request on-chain

Servers lock inference requests, once again, by sending a transaction to a smart contract.

This transaction includes the peer to peer address the client should send the prompt to.

4. Client sends prompt to server off-chain

The client then sends the prompt to the provided peer to peer address.

This should be encrypted - there may be encryption built into the peer to peer library or protocol we end up using, otherwise we may have client and server sharing public keys during the request submitting and request locking process.

5. Server streams response to client off-chain

After receiving the prompt, the server streams the response to the peer to peer address previously provided by the client.

This response, similarly to the prompt, should also be encrypted.

6. Server claims payment on-chain

When a server finishes responding to an inference request, it can charge the corresponding cost from the amount previously staked by the client, again by submitting a transaction to a smart contract.

After locking an inference request, servers have a 1 hour limit to claim payment. If they don't, the request is cancelled and the stake is fully returned to the user.

This should be a rare case, as there is no incentive for servers to not claim payment on a request, malicious or not.

This limit does handle the case where a benevolent server fails to respond to a request. The server may still get flagged by the client regardless, and likely before this time limit is reached, affecting its reputation.

Reputation system

Malicious behavior is disincentivized through a reputation system.

Monetization

We can introduce a fee (maybe around 10%) on the cost of each inference through the smart contract at some point in the future.

Question: should this be charged from the client's stake in addition to the server's charge, or deducted from the server's charge? Maybe the latter, as servers are the ones profiting so they should be the ones deducting.

Storage

TODO

Could be IPFS/filecoin or arweave

Website

Chainference will have a centralized website with:

  • List of active servers and their details
  • List of models and their details such as pricing and server availability

Solana facts

  • Fees are cheap, with $0.01 being enough for around 8k transactions at the base fee with a SOL value of $250. Prioritization fees are almost negligible at an average of ~1 lamport (1e-9 SOL).
  • Storage is expensive, with a cost of ~7 SOL per MB ($1750 at $250 SOL price). This value is however fully recoverable on deletion.

Thoughts on privacy

Anonymity

Prompts and responses are only visible to each client and server engaging in a transaction.

However, anonymity could be enhanced by using disposable accounts to submit prompts with.

To avoid tracing back to the real wallet, a centralized place could be creating these disposable wallets for everyone, perhaps the chainference smart contract itself.

However, servers can still associate prompts to the address they are sending the response to.

So the possibility of users creating temporary addresses to receive responses at should be investigated - some sort of proxying.

Encryption

Future improvements in encrypted inference may allow us to have servers running inference on prompts without being able to see its plain text contents, and perhaps also for the response.

About

LLM inference on the solana chain.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Chainference

Decentralized inference on the Solana chain.

As a server, you can monetize your hardware by running paid AI inference for users.

As a user, you get access to open AI models without restrictions and for cheaper than on centralized services.

Cheaper because unrestricted competition drives prices down to healthy profit margins.

Getting started

See each subfolder's readme for instructions on the different parts of the project.

How it works

  1. Servers publish availability on-chain
  2. Clients publish inference requests on-chain, staking maximum desired cost
  3. A server locks the inference request on-chain
  4. Client sends prompt to server off-chain
  5. Server streams response to client off-chain
  6. Server claims payment on-chain, with remainder returned to client

1. Servers publish availability on-chain

Servers publish availability by sending a transaction to a smart contract, including a list of models available for inference (hugging face IDs) and price per million output tokens for each.

2. Clients publish inference requests on-chain

Clients publish inference requests on chain also by sending a transaction to a smart contract.

They don't send requests directly to servers because the maximum cost must be staked on-chain to ensure payment.

An inference request includes:

  • Filtering criteria for server
  • Desired model
  • SOL stake of maximum cost
  • Creation timestamp

3. A server locks the inference request on-chain

Servers lock inference requests, once again, by sending a transaction to a smart contract.

This transaction includes the peer to peer address the client should send the prompt to.

4. Client sends prompt to server off-chain

The client then sends the prompt to the provided peer to peer address.

This should be encrypted - there may be encryption built into the peer to peer library or protocol we end up using, otherwise we may have client and server sharing public keys during the request submitting and request locking process.

5. Server streams response to client off-chain

After receiving the prompt, the server streams the response to the peer to peer address previously provided by the client.

This response, similarly to the prompt, should also be encrypted.

6. Server claims payment on-chain

When a server finishes responding to an inference request, it can charge the corresponding cost from the amount previously staked by the client, again by submitting a transaction to a smart contract.

After locking an inference request, servers have a 1 hour limit to claim payment. If they don't, the request is cancelled and the stake is fully returned to the user.

This should be a rare case, as there is no incentive for servers to not claim payment on a request, malicious or not.

This limit does handle the case where a benevolent server fails to respond to a request. The server may still get flagged by the client regardless, and likely before this time limit is reached, affecting its reputation.

Reputation system

Malicious behavior is disincentivized through a reputation system.

Monetization

We can introduce a fee (maybe around 10%) on the cost of each inference through the smart contract at some point in the future.

Question: should this be charged from the client's stake in addition to the server's charge, or deducted from the server's charge? Maybe the latter, as servers are the ones profiting so they should be the ones deducting.

Storage

TODO

Could be IPFS/filecoin or arweave

Website

Chainference will have a centralized website with:

  • List of active servers and their details
  • List of models and their details such as pricing and server availability

Solana facts

  • Fees are cheap, with $0.01 being enough for around 8k transactions at the base fee with a SOL value of $250. Prioritization fees are almost negligible at an average of ~1 lamport (1e-9 SOL).
  • Storage is expensive, with a cost of ~7 SOL per MB ($1750 at $250 SOL price). This value is however fully recoverable on deletion.

Thoughts on privacy

Anonymity

Prompts and responses are only visible to each client and server engaging in a transaction.

However, anonymity could be enhanced by using disposable accounts to submit prompts with.

To avoid tracing back to the real wallet, a centralized place could be creating these disposable wallets for everyone, perhaps the chainference smart contract itself.

However, servers can still associate prompts to the address they are sending the response to.

So the possibility of users creating temporary addresses to receive responses at should be investigated - some sort of proxying.

Encryption

Future improvements in encrypted inference may allow us to have servers running inference on prompts without being able to see its plain text contents, and perhaps also for the response.

About

LLM inference on the solana chain.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Chainference

Decentralized inference on the Solana chain.

As a server, you can monetize your hardware by running paid AI inference for users.

As a user, you get access to open AI models without restrictions and for cheaper than on centralized services.

Cheaper because unrestricted competition drives prices down to healthy profit margins.

Getting started

See each subfolder's readme for instructions on the different parts of the project.

How it works

  1. Servers publish availability on-chain
  2. Clients publish inference requests on-chain, staking maximum desired cost
  3. A server locks the inference request on-chain
  4. Client sends prompt to server off-chain
  5. Server streams response to client off-chain
  6. Server claims payment on-chain, with remainder returned to client

1. Servers publish availability on-chain

Servers publish availability by sending a transaction to a smart contract, including a list of models available for inference (hugging face IDs) and price per million output tokens for each.

2. Clients publish inference requests on-chain

Clients publish inference requests on chain also by sending a transaction to a smart contract.

They don't send requests directly to servers because the maximum cost must be staked on-chain to ensure payment.

An inference request includes:

  • Filtering criteria for server
  • Desired model
  • SOL stake of maximum cost
  • Creation timestamp

3. A server locks the inference request on-chain

Servers lock inference requests, once again, by sending a transaction to a smart contract.

This transaction includes the peer to peer address the client should send the prompt to.

4. Client sends prompt to server off-chain

The client then sends the prompt to the provided peer to peer address.

This should be encrypted - there may be encryption built into the peer to peer library or protocol we end up using, otherwise we may have client and server sharing public keys during the request submitting and request locking process.

5. Server streams response to client off-chain

After receiving the prompt, the server streams the response to the peer to peer address previously provided by the client.

This response, similarly to the prompt, should also be encrypted.

6. Server claims payment on-chain

When a server finishes responding to an inference request, it can charge the corresponding cost from the amount previously staked by the client, again by submitting a transaction to a smart contract.

After locking an inference request, servers have a 1 hour limit to claim payment. If they don't, the request is cancelled and the stake is fully returned to the user.

This should be a rare case, as there is no incentive for servers to not claim payment on a request, malicious or not.

This limit does handle the case where a benevolent server fails to respond to a request. The server may still get flagged by the client regardless, and likely before this time limit is reached, affecting its reputation.

Reputation system

Malicious behavior is disincentivized through a reputation system.

Monetization

We can introduce a fee (maybe around 10%) on the cost of each inference through the smart contract at some point in the future.

Question: should this be charged from the client's stake in addition to the server's charge, or deducted from the server's charge? Maybe the latter, as servers are the ones profiting so they should be the ones deducting.

Storage

TODO

Could be IPFS/filecoin or arweave

Website

Chainference will have a centralized website with:

  • List of active servers and their details
  • List of models and their details such as pricing and server availability

Solana facts

  • Fees are cheap, with $0.01 being enough for around 8k transactions at the base fee with a SOL value of $250. Prioritization fees are almost negligible at an average of ~1 lamport (1e-9 SOL).
  • Storage is expensive, with a cost of ~7 SOL per MB ($1750 at $250 SOL price). This value is however fully recoverable on deletion.

Thoughts on privacy

Anonymity

Prompts and responses are only visible to each client and server engaging in a transaction.

However, anonymity could be enhanced by using disposable accounts to submit prompts with.

To avoid tracing back to the real wallet, a centralized place could be creating these disposable wallets for everyone, perhaps the chainference smart contract itself.

However, servers can still associate prompts to the address they are sending the response to.

So the possibility of users creating temporary addresses to receive responses at should be investigated - some sort of proxying.

Encryption

Future improvements in encrypted inference may allow us to have servers running inference on prompts without being able to see its plain text contents, and perhaps also for the response.

About

LLM inference on the solana chain.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Chainference

Decentralized inference on the Solana chain.

As a server, you can monetize your hardware by running paid AI inference for users.

As a user, you get access to open AI models without restrictions and for cheaper than on centralized services.

Cheaper because unrestricted competition drives prices down to healthy profit margins.

Getting started

See each subfolder's readme for instructions on the different parts of the project.

How it works

  1. Servers publish availability on-chain
  2. Clients publish inference requests on-chain, staking maximum desired cost
  3. A server locks the inference request on-chain
  4. Client sends prompt to server off-chain
  5. Server streams response to client off-chain
  6. Server claims payment on-chain, with remainder returned to client

1. Servers publish availability on-chain

Servers publish availability by sending a transaction to a smart contract, including a list of models available for inference (hugging face IDs) and price per million output tokens for each.

2. Clients publish inference requests on-chain

Clients publish inference requests on chain also by sending a transaction to a smart contract.

They don't send requests directly to servers because the maximum cost must be staked on-chain to ensure payment.

An inference request includes:

  • Filtering criteria for server
  • Desired model
  • SOL stake of maximum cost
  • Creation timestamp

3. A server locks the inference request on-chain

Servers lock inference requests, once again, by sending a transaction to a smart contract.

This transaction includes the peer to peer address the client should send the prompt to.

4. Client sends prompt to server off-chain

The client then sends the prompt to the provided peer to peer address.

This should be encrypted - there may be encryption built into the peer to peer library or protocol we end up using, otherwise we may have client and server sharing public keys during the request submitting and request locking process.

5. Server streams response to client off-chain

After receiving the prompt, the server streams the response to the peer to peer address previously provided by the client.

This response, similarly to the prompt, should also be encrypted.

6. Server claims payment on-chain

When a server finishes responding to an inference request, it can charge the corresponding cost from the amount previously staked by the client, again by submitting a transaction to a smart contract.

After locking an inference request, servers have a 1 hour limit to claim payment. If they don't, the request is cancelled and the stake is fully returned to the user.

This should be a rare case, as there is no incentive for servers to not claim payment on a request, malicious or not.

This limit does handle the case where a benevolent server fails to respond to a request. The server may still get flagged by the client regardless, and likely before this time limit is reached, affecting its reputation.

Reputation system

Malicious behavior is disincentivized through a reputation system.

Monetization

We can introduce a fee (maybe around 10%) on the cost of each inference through the smart contract at some point in the future.

Question: should this be charged from the client's stake in addition to the server's charge, or deducted from the server's charge? Maybe the latter, as servers are the ones profiting so they should be the ones deducting.

Storage

TODO

Could be IPFS/filecoin or arweave

Website

Chainference will have a centralized website with:

  • List of active servers and their details
  • List of models and their details such as pricing and server availability

Solana facts

  • Fees are cheap, with $0.01 being enough for around 8k transactions at the base fee with a SOL value of $250. Prioritization fees are almost negligible at an average of ~1 lamport (1e-9 SOL).
  • Storage is expensive, with a cost of ~7 SOL per MB ($1750 at $250 SOL price). This value is however fully recoverable on deletion.

Thoughts on privacy

Anonymity

Prompts and responses are only visible to each client and server engaging in a transaction.

However, anonymity could be enhanced by using disposable accounts to submit prompts with.

To avoid tracing back to the real wallet, a centralized place could be creating these disposable wallets for everyone, perhaps the chainference smart contract itself.

However, servers can still associate prompts to the address they are sending the response to.

So the possibility of users creating temporary addresses to receive responses at should be investigated - some sort of proxying.

Encryption

Future improvements in encrypted inference may allow us to have servers running inference on prompts without being able to see its plain text contents, and perhaps also for the response.

About

LLM inference on the solana chain.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Chainference

Decentralized inference on the Solana chain.

As a server, you can monetize your hardware by running paid AI inference for users.

As a user, you get access to open AI models without restrictions and for cheaper than on centralized services.

Cheaper because unrestricted competition drives prices down to healthy profit margins.

Getting started

See each subfolder's readme for instructions on the different parts of the project.

How it works

  1. Servers publish availability on-chain
  2. Clients publish inference requests on-chain, staking maximum desired cost
  3. A server locks the inference request on-chain
  4. Client sends prompt to server off-chain
  5. Server streams response to client off-chain
  6. Server claims payment on-chain, with remainder returned to client

1. Servers publish availability on-chain

Servers publish availability by sending a transaction to a smart contract, including a list of models available for inference (hugging face IDs) and price per million output tokens for each.

2. Clients publish inference requests on-chain

Clients publish inference requests on chain also by sending a transaction to a smart contract.

They don't send requests directly to servers because the maximum cost must be staked on-chain to ensure payment.

An inference request includes:

  • Filtering criteria for server
  • Desired model
  • SOL stake of maximum cost
  • Creation timestamp

3. A server locks the inference request on-chain

Servers lock inference requests, once again, by sending a transaction to a smart contract.

This transaction includes the peer to peer address the client should send the prompt to.

4. Client sends prompt to server off-chain

The client then sends the prompt to the provided peer to peer address.

This should be encrypted - there may be encryption built into the peer to peer library or protocol we end up using, otherwise we may have client and server sharing public keys during the request submitting and request locking process.

5. Server streams response to client off-chain

After receiving the prompt, the server streams the response to the peer to peer address previously provided by the client.

This response, similarly to the prompt, should also be encrypted.

6. Server claims payment on-chain

When a server finishes responding to an inference request, it can charge the corresponding cost from the amount previously staked by the client, again by submitting a transaction to a smart contract.

After locking an inference request, servers have a 1 hour limit to claim payment. If they don't, the request is cancelled and the stake is fully returned to the user.

This should be a rare case, as there is no incentive for servers to not claim payment on a request, malicious or not.

This limit does handle the case where a benevolent server fails to respond to a request. The server may still get flagged by the client regardless, and likely before this time limit is reached, affecting its reputation.

Reputation system

Malicious behavior is disincentivized through a reputation system.

Monetization

We can introduce a fee (maybe around 10%) on the cost of each inference through the smart contract at some point in the future.

Question: should this be charged from the client's stake in addition to the server's charge, or deducted from the server's charge? Maybe the latter, as servers are the ones profiting so they should be the ones deducting.

Storage

TODO

Could be IPFS/filecoin or arweave

Website

Chainference will have a centralized website with:

  • List of active servers and their details
  • List of models and their details such as pricing and server availability

Solana facts

  • Fees are cheap, with $0.01 being enough for around 8k transactions at the base fee with a SOL value of $250. Prioritization fees are almost negligible at an average of ~1 lamport (1e-9 SOL).
  • Storage is expensive, with a cost of ~7 SOL per MB ($1750 at $250 SOL price). This value is however fully recoverable on deletion.

Thoughts on privacy

Anonymity

Prompts and responses are only visible to each client and server engaging in a transaction.

However, anonymity could be enhanced by using disposable accounts to submit prompts with.

To avoid tracing back to the real wallet, a centralized place could be creating these disposable wallets for everyone, perhaps the chainference smart contract itself.

However, servers can still associate prompts to the address they are sending the response to.

So the possibility of users creating temporary addresses to receive responses at should be investigated - some sort of proxying.

Encryption

Future improvements in encrypted inference may allow us to have servers running inference on prompts without being able to see its plain text contents, and perhaps also for the response.

About

LLM inference on the solana chain.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Chainference

Decentralized inference on the Solana chain.

As a server, you can monetize your hardware by running paid AI inference for users.

As a user, you get access to open AI models without restrictions and for cheaper than on centralized services.

Cheaper because unrestricted competition drives prices down to healthy profit margins.

Getting started

See each subfolder's readme for instructions on the different parts of the project.

How it works

  1. Servers publish availability on-chain
  2. Clients publish inference requests on-chain, staking maximum desired cost
  3. A server locks the inference request on-chain
  4. Client sends prompt to server off-chain
  5. Server streams response to client off-chain
  6. Server claims payment on-chain, with remainder returned to client

1. Servers publish availability on-chain

Servers publish availability by sending a transaction to a smart contract, including a list of models available for inference (hugging face IDs) and price per million output tokens for each.

2. Clients publish inference requests on-chain

Clients publish inference requests on chain also by sending a transaction to a smart contract.

They don't send requests directly to servers because the maximum cost must be staked on-chain to ensure payment.

An inference request includes:

  • Filtering criteria for server
  • Desired model
  • SOL stake of maximum cost
  • Creation timestamp

3. A server locks the inference request on-chain

Servers lock inference requests, once again, by sending a transaction to a smart contract.

This transaction includes the peer to peer address the client should send the prompt to.

4. Client sends prompt to server off-chain

The client then sends the prompt to the provided peer to peer address.

This should be encrypted - there may be encryption built into the peer to peer library or protocol we end up using, otherwise we may have client and server sharing public keys during the request submitting and request locking process.

5. Server streams response to client off-chain

After receiving the prompt, the server streams the response to the peer to peer address previously provided by the client.

This response, similarly to the prompt, should also be encrypted.

6. Server claims payment on-chain

When a server finishes responding to an inference request, it can charge the corresponding cost from the amount previously staked by the client, again by submitting a transaction to a smart contract.

After locking an inference request, servers have a 1 hour limit to claim payment. If they don't, the request is cancelled and the stake is fully returned to the user.

This should be a rare case, as there is no incentive for servers to not claim payment on a request, malicious or not.

This limit does handle the case where a benevolent server fails to respond to a request. The server may still get flagged by the client regardless, and likely before this time limit is reached, affecting its reputation.

Reputation system

Malicious behavior is disincentivized through a reputation system.

Monetization

We can introduce a fee (maybe around 10%) on the cost of each inference through the smart contract at some point in the future.

Question: should this be charged from the client's stake in addition to the server's charge, or deducted from the server's charge? Maybe the latter, as servers are the ones profiting so they should be the ones deducting.

Storage

TODO

Could be IPFS/filecoin or arweave

Website

Chainference will have a centralized website with:

  • List of active servers and their details
  • List of models and their details such as pricing and server availability

Solana facts

  • Fees are cheap, with $0.01 being enough for around 8k transactions at the base fee with a SOL value of $250. Prioritization fees are almost negligible at an average of ~1 lamport (1e-9 SOL).
  • Storage is expensive, with a cost of ~7 SOL per MB ($1750 at $250 SOL price). This value is however fully recoverable on deletion.

Thoughts on privacy

Anonymity

Prompts and responses are only visible to each client and server engaging in a transaction.

However, anonymity could be enhanced by using disposable accounts to submit prompts with.

To avoid tracing back to the real wallet, a centralized place could be creating these disposable wallets for everyone, perhaps the chainference smart contract itself.

However, servers can still associate prompts to the address they are sending the response to.

So the possibility of users creating temporary addresses to receive responses at should be investigated - some sort of proxying.

Encryption

Future improvements in encrypted inference may allow us to have servers running inference on prompts without being able to see its plain text contents, and perhaps also for the response.

About

LLM inference on the solana chain.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages