From 94317ad0fb33e6cc0e416e3856de98a487c7c6a3 Mon Sep 17 00:00:00 2001 From: Aaron Surty Date: Fri, 18 Mar 2022 09:50:12 -0400 Subject: [PATCH 1/5] template setup private-data-backup RFC (#58) --- docs/RFCs/private-data-backup.md | 37 ++++++++++++++++++++++++++++++++ 1 file changed, 37 insertions(+) create mode 100644 docs/RFCs/private-data-backup.md diff --git a/docs/RFCs/private-data-backup.md b/docs/RFCs/private-data-backup.md new file mode 100644 index 0000000..80ef7e4 --- /dev/null +++ b/docs/RFCs/private-data-backup.md @@ -0,0 +1,37 @@ +- Feature Name: private-data-backup +- Start Date: 2022-03-18 +- RFC PR: @TODO +- Functionland Issue: https://github.com/functionland/docs/issues/58 + + +# Summary +[summary]: #summary + +@TODO + +# Motivation +[motivation]: #motivation + +# Guide-level explanation +[guide-level-explanation]: #guide-level-explanation + + +# Reference-level explanation +[reference-level-explanation]: #reference-level-explanation + +# Drawbacks +[drawbacks]: #drawbacks + + +# Rationale and alternatives +[rationale-and-alternatives]: #rationale-and-alternatives + +# Prior art +[prior-art]: #prior-art + +# Unresolved questions +[unresolved-questions]: #unresolved-questions + +# Future possibilities +[future-possibilities]: #future-possibilities + From ebcc76ac41df2ecd876cbf16d2c94399c1bc6dd8 Mon Sep 17 00:00:00 2001 From: Aaron Surty Date: Fri, 18 Mar 2022 10:10:01 -0400 Subject: [PATCH 2/5] first draft of personal data backup RFC (#58) --- docs/RFCs/:w | 248 ++++++++++++++++++++++++++++++ docs/RFCs/personal-data-backup.md | 246 +++++++++++++++++++++++++++++ docs/RFCs/private-data-backup.md | 37 ----- 3 files changed, 494 insertions(+), 37 deletions(-) create mode 100644 docs/RFCs/:w create mode 100644 docs/RFCs/personal-data-backup.md delete mode 100644 docs/RFCs/private-data-backup.md diff --git a/docs/RFCs/:w b/docs/RFCs/:w new file mode 100644 index 0000000..b2496b6 --- /dev/null +++ b/docs/RFCs/:w @@ -0,0 +1,248 @@ +- Feature Name: personal-data-backup +- Start Date: 2022-03-18 +- RFC PR: [functionland/docs/pull/61](https://github.com/functionland/docs/pull/61) +- Functionland Issue: [functionland/docs/issues/58](https://github.com/functionland/docs/issues/58) +- Status: Draft +- Authors: [Aaron Surty](https://github.com/gitaaron), [Farhoud](https://github.com/farhoud) +- Reviewers: @TODO + +# Summary +[summary]: #summary + +This RFC covers how a pool of BOXes can work collaboratively together to improve data reliability. + +# Motivation +[motivation]: #motivation + +A person owns several BOXes and wants a portion of their data replicated across each BOX so that if one of the BOXes malfunctions they should not lose any of their data. + +The following scenarios are handled: + + * adding a new BOX to the pool + + * removing a BOX from the pool + + * adding a hard drive to a BOX + + * removing a hard drive from a BOX + + * a severe network outage occurs severing a region of BOXes from another region + + * data becomes corrupted on a BOX + + * updating the same data set in real time + + +# Guide-level explanation +[guide-level-explanation]: #guide-level-explanation + +## Terminology +[terminology]: #terminology + +| Name | Definition | +|----------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------| +| snapshot | the entire collection of data that might be stored in chunks across several keepers but can be rebuilt to represent the entire file system being backed up | +| keeper | a BOX process responsible for storing a portion of a snapshot and sharing the burden of recreating an entire snapshot | +| BOX | an OS that each keeper runs on | +| author | a BOX where the data set was created or written to last | +| replication factor | how many keepers a chunk of data is stored on; a greater replication factor means greater reliability | +| pool | a group of keepers collaboratively working together to store a snapshot | +| region | a subgroup of BOXes within a pool by grouped by geographic proximity to each other | +| data set | a file, directory, or a discrete piece of data sitting in a database | + + +## Pre-conditions +[pre-conditions]: #pre-conditions + +* each BOX is already provisioned with the necessary configuration info in order for the keeper to fully operate +* each keeper in a pool can be trusted to not operate maliciously +* the type of file system that each BOX is backing up is the same + +## Limitations +[limitations]: #limitations + +The following limitations may be encountered while operating a pool: +* file size +* number of files in a directory +* snapshot size +* number of keepers in a pool + +## Configuration + +Configuration data for each node can be split into local and shared. + +### Local + + * local BOX address + * local public/private key + +### Shared + + * remote BOX addresses of participants in the pool + * shared secret + * minimum acceptable replication factor + * warning time (how much time should be given for a warning to be sent out before a limitation occurs) + * has a global default as well as an override for each custom limitation + + +## Local Filesystem +If space permits, the entire contents of a snapshot may be stored on an author's local FS. If the author runs out of space then contents must be sharded across several BOXes. + +If the actual replication factor drops below the minimum acceptable threshold a failure event is dispatched. + +## Conflict Resolution +[conflict-resolution]: #conflict-resolution +A conflict may arise between keepers when a snapshat goes out of sync. This could occur either due to [real time updates](#real-time-updates) or [disk corruption](#disk-corruption). + +### Real-time Updates +[real-time-updates]: #real-time-updates + * more than one person is editing the same data set on several BOXes in a pool at the same time + +### Corruption +[coruption]: #corruption + * a disk becomes corrupt on a BOX + +In either case, the disputing keepers should take the appropriate steps to resolve the conflict. If an appropriate back-out strategy can not be achieved an event should be raised. + +## Events +The following events should be dispatched for an administrative UI. + + * limitation encountered + + * network disruption + + * keeper health + * memory + * disk I/O + * CPU usage + * disk corruption + + * unresolved conflict + + * unacceptable replication factor + +### Event Types + * normal + * warning + * failure + +### Warnings + +Events dispatched before a failure occurs based on a forecasting heuristic to determine how quickly a limit will be reached. + +### Event Frequency +Normal events are dispatched periodically (based on a config param) for historical reporting and warnings|failures are dispatched immediately. + +# Reference-level Explanation +[reference-level-explanation]: #reference-level-explanation + +## Network Architecture + +A peer-peer architecture is used over master/slave so that if a single keeper goes down the rest of the pool will still be able to operate normally. + + * any shared config data is stored on each BOX + + * any shared state required for the retrieval of a data set is store on each BOX + + * no central servers are used for routing + + +## Availability + +@TODO + +Describe some heuristics that can be used to achieve greater availability (ie/ storing entire contents of data set on author) + +### Conflict Resolution + +@TODO + +For detecting / handling conflicts due to data sets going out of sync from real-time updates + +There are different conflict resolution strategies available. + +Both strategies might be used by a single keeper depending on the type of data set or an override. + +### File Integrity Monitoring + +@TODO + +For detecting / handling data set corruption. + +Comparing the contents of chunks on disk with a source of truth. + +The source of truth is only be updated from change events. + + +## Event Dispatcher + +@TODO + +## Data Set Retrieval + +@TODO + +## Regions + +@TODO + +A best effort approach to storing an entire snapshot in a single region so that if a region is severed the pool will still be able to recover the entire snapshot. + +## Network Stack + +@TODO + +Defines how peers will communicate with each other. + +# Drawbacks +[drawbacks]: #drawbacks +Putting the responsibility of data reliability on BOX owners means there is potential for a BOX owner to make a mistake and permanently lose their data. + +# Rationale and alternatives +[rationale-and-alternatives]: #rationale-and-alternatives +An alternative could be to use paid services (cloud storage providers) with their own SLAs to take on the responsibility of data reliability. + +If participating in a decentralized storage network (DSN), the BOX owner could also purchase a mining component to offset their cost. + +There are currently a few drawbacks with this: + + * becoming a storage miner requires a significant upfront investment to cover hardware and staking costs + + * a private pool will always be more efficient since keepers will not have to worry about the overhead of trusting each other + +These options are not mutually exclusive. Offering both options (free and paid) could provide the greatest freedom/flexibility for BOX owners. + +# Prior art +[prior-art]: #prior-art + + * [IPFS cluster](https://cluster.ipfs.io/) + +# Unresolved questions +[unresolved-questions]: #unresolved-questions + +How is a data set reconstructed from chunks? + +How are chunks found on the network? (content discovery) + +What is an ideal replication factor? + +How can contents of entire filesystem be efficiently compared with snapshot (aka source of truth)? + +How can reliability be measured? Markov models? + +How can limits be calculated? + +How can NAT hole punching work in a pnet without any relays? + +Should we consider using a VPN or other alternative such as Tor over libp2p? + +How will bootstrapping work for BOXes not on same LAN? + +Which components can be re-used for data sharing? + +Does data compression need to be taken into account? + +# Future possibilities +[future-possibilities]: #future-possibilities + +Storing a history of the snapshot so an owner can go back in time and recover a data set from a previous state. diff --git a/docs/RFCs/personal-data-backup.md b/docs/RFCs/personal-data-backup.md new file mode 100644 index 0000000..73f326b --- /dev/null +++ b/docs/RFCs/personal-data-backup.md @@ -0,0 +1,246 @@ +- Feature Name: personal-data-backup +- Start Date: 2022-03-18 +- RFC PR: [functionland/docs/pull/61](https://github.com/functionland/docs/pull/61) +- Functionland Issue: [functionland/docs/issues/58](https://github.com/functionland/docs/issues/58) +- Status: Draft +- Authors: [Aaron Surty](https://github.com/gitaaron), [Farhoud](https://github.com/farhoud) +- Reviewers: @TODO + +# Summary +[summary]: #summary + +This RFC covers how a pool of BOXes can work collaboratively together to improve data reliability. + +# Motivation +[motivation]: #motivation + +A person owns several BOXes and wants a portion of their data replicated across each BOX so that if one of the BOXes malfunctions they should not lose any of their data. + +The following scenarios are handled: + + * adding a new BOX to the pool + + * removing a BOX from the pool + + * adding a hard drive to a BOX + + * removing a hard drive from a BOX + + * a severe network outage occurs severing a region of BOXes from another region + + * data becomes corrupted on a BOX + + * updating the same data set in real time + + +# Guide-level explanation +[guide-level-explanation]: #guide-level-explanation + +## Terminology +[terminology]: #terminology + +| Name | Definition | +|----------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------| +| snapshot | the entire collection of data that might be stored in chunks across several keepers but can be rebuilt to represent the entire file system being backed up | +| keeper | a BOX process responsible for storing a portion of a snapshot and sharing the burden of recreating an entire snapshot | +| BOX | an OS that each keeper runs on | +| author | a BOX where the data set was created or written to last | +| replication factor | how many keepers a chunk of data is stored on; a greater replication factor means greater reliability | +| pool | a group of keepers collaboratively working together to store a snapshot | +| region | a subgroup of BOXes within a pool by grouped by geographic proximity to each other | +| data set | a file, directory, or a discrete piece of data sitting in a database | + + +## Pre-conditions +[pre-conditions]: #pre-conditions + +* each BOX is already provisioned with the necessary configuration info in order for the keeper to fully operate +* each keeper in a pool can be trusted to not operate maliciously +* the type of file system that each BOX is backing up is the same + +## Limitations +[limitations]: #limitations + +The following limitations may be encountered while operating a pool: +* file size +* number of files in a directory +* snapshot size +* number of keepers in a pool + +## Configuration + +Configuration data for each node can be split into local and shared. + +### Local + + * local BOX address + * local public/private key + +### Shared + + * remote BOX addresses of participants in the pool + * shared secret + * minimum acceptable replication factor + * normal event frequency + * warning time + * how much time should be given for a warning to be sent out before a imminent limitation is encountered and a system failure occurs + * has a global default as well as an override for each custom limitation + +## Conflict Resolution +[conflict-resolution]: #conflict-resolution +A conflict may arise between keepers when a snapshat goes out of sync. This could occur either due to [real time updates](#real-time-updates) or [disk corruption](#disk-corruption). + +### Real-time Updates +[real-time-updates]: #real-time-updates + * more than one person is editing the same data set on several BOXes in a pool at the same time + +### Disk Corruption +[disk-corruption]: #disk-corruption + * a disk becomes corrupt on a BOX + +In either case, the disputing keepers will take the appropriate steps to resolve the conflict. If an appropriate back-out strategy can not be achieved, an event is raised. + +## Events +The following events should be dispatched for an administrative UI. + + * limitation imminent + + * limitation encountered + + * keeper health + * memory + * disk I/O + * CPU usage + * disk corruption + + * unresolved conflict + + * unacceptable replication factor + + * keeper added / removed + + * network disruption + + +### Event Types + * normal + * warning + * failure + +### Warnings + +Events dispatched before a failure occurs based on a forecasting heuristic to determine how quickly a limit will be reached. + +### Event Frequency +Normal events are dispatched periodically (based on a config param) for historical reporting and warnings|failures are dispatched immediately. + +## Regions + +If multiple regions are set up, each region contains the entire contents of a snapshot. If a region is severed from the pool, it will still be able to recover the entire snapshot. + +# Reference-level Explanation +[reference-level-explanation]: #reference-level-explanation + +## Network Architecture + +A peer-peer architecture is used over master/slave so that if a single keeper goes down the rest of the pool will still be able to operate normally. + + * any shared config data is stored on each BOX + + * any shared state required for the retrieval of a data set is stored on each BOX + + * no central servers are used for routing + + +## Heuristics + +Some heuristics can be used to achieve greater availability / load times and minimize bandwidth. + +### Local File System First + +If space permits, the entire contents of a snapshot may be stored on an author's local file system. If the author runs out of space then contents must be sharded across several BOXes. + +@TODO - fill out other heuristics + + +### Conflict Resolution + +For detecting / handling conflicts due to data sets going out of sync from real-time updates + +There are different conflict resolution strategies available. + +Both strategies might be used by a single keeper depending on the type of data set or an override. + +### File Integrity Monitoring + +For detecting / handling data set corruption. + +Comparing the contents of chunks on disk with a source of truth. + +The source of truth is only updated from change events. + +## Event Dispatcher + +@TODO + +## Data Set Retrieval + +@TODO + + +## Network Stack + +@TODO + +# Drawbacks +[drawbacks]: #drawbacks +Putting the responsibility of data reliability on BOX owners means there is potential for a BOX owner to make a mistake and permanently lose their data. + +# Rationale and alternatives +[rationale-and-alternatives]: #rationale-and-alternatives +An alternative could be to use paid services (cloud storage providers) with their own SLAs to take on the responsibility of data reliability. + +If participating in a decentralized storage network (DSN), the BOX owner could also purchase a mining component to offset their cost. + +There are currently a few drawbacks with this: + + * becoming a storage miner requires a significant upfront investment to cover hardware and staking costs + + * a private pool will always be more efficient since keepers will not have to worry about the overhead of trusting each other + +These options are not mutually exclusive. Offering both options (free and paid) could provide the greatest freedom/flexibility for BOX owners. + +# Prior art +[prior-art]: #prior-art + + * [IPFS cluster](https://cluster.ipfs.io/) + +# Unresolved questions +[unresolved-questions]: #unresolved-questions + +How is a data set reconstructed from chunks? + +How are chunks found on the network? (content discovery) + +What is an ideal replication factor? + +How can contents of an entire filesystem be efficiently compared with snapshot (aka source of truth)? + +How can reliability be measured? Markov models? + +How can system limits be calculated? + +How can NAT hole punching work in a pnet without any relays? + +Should we consider using a VPN or other alternatives such as Tor over libp2p? + +How will bootstrapping work for BOXes not on same LAN? + +Which components can be re-used for data sharing? + +Does data compression need to be taken into account? + +# Future possibilities +[future-possibilities]: #future-possibilities + +Storing a history of the snapshot so an owner can go back in time and recover a data set from a previous state. diff --git a/docs/RFCs/private-data-backup.md b/docs/RFCs/private-data-backup.md deleted file mode 100644 index 80ef7e4..0000000 --- a/docs/RFCs/private-data-backup.md +++ /dev/null @@ -1,37 +0,0 @@ -- Feature Name: private-data-backup -- Start Date: 2022-03-18 -- RFC PR: @TODO -- Functionland Issue: https://github.com/functionland/docs/issues/58 - - -# Summary -[summary]: #summary - -@TODO - -# Motivation -[motivation]: #motivation - -# Guide-level explanation -[guide-level-explanation]: #guide-level-explanation - - -# Reference-level explanation -[reference-level-explanation]: #reference-level-explanation - -# Drawbacks -[drawbacks]: #drawbacks - - -# Rationale and alternatives -[rationale-and-alternatives]: #rationale-and-alternatives - -# Prior art -[prior-art]: #prior-art - -# Unresolved questions -[unresolved-questions]: #unresolved-questions - -# Future possibilities -[future-possibilities]: #future-possibilities - From 2c8cdc5b5c6afd1a0ca4a2d12558077bceaf332a Mon Sep 17 00:00:00 2001 From: Aaron Surty Date: Fri, 18 Mar 2022 23:14:52 -0400 Subject: [PATCH 3/5] removing erroneously added file (#58) --- docs/RFCs/:w | 248 --------------------------------------------------- 1 file changed, 248 deletions(-) delete mode 100644 docs/RFCs/:w diff --git a/docs/RFCs/:w b/docs/RFCs/:w deleted file mode 100644 index b2496b6..0000000 --- a/docs/RFCs/:w +++ /dev/null @@ -1,248 +0,0 @@ -- Feature Name: personal-data-backup -- Start Date: 2022-03-18 -- RFC PR: [functionland/docs/pull/61](https://github.com/functionland/docs/pull/61) -- Functionland Issue: [functionland/docs/issues/58](https://github.com/functionland/docs/issues/58) -- Status: Draft -- Authors: [Aaron Surty](https://github.com/gitaaron), [Farhoud](https://github.com/farhoud) -- Reviewers: @TODO - -# Summary -[summary]: #summary - -This RFC covers how a pool of BOXes can work collaboratively together to improve data reliability. - -# Motivation -[motivation]: #motivation - -A person owns several BOXes and wants a portion of their data replicated across each BOX so that if one of the BOXes malfunctions they should not lose any of their data. - -The following scenarios are handled: - - * adding a new BOX to the pool - - * removing a BOX from the pool - - * adding a hard drive to a BOX - - * removing a hard drive from a BOX - - * a severe network outage occurs severing a region of BOXes from another region - - * data becomes corrupted on a BOX - - * updating the same data set in real time - - -# Guide-level explanation -[guide-level-explanation]: #guide-level-explanation - -## Terminology -[terminology]: #terminology - -| Name | Definition | -|----------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------| -| snapshot | the entire collection of data that might be stored in chunks across several keepers but can be rebuilt to represent the entire file system being backed up | -| keeper | a BOX process responsible for storing a portion of a snapshot and sharing the burden of recreating an entire snapshot | -| BOX | an OS that each keeper runs on | -| author | a BOX where the data set was created or written to last | -| replication factor | how many keepers a chunk of data is stored on; a greater replication factor means greater reliability | -| pool | a group of keepers collaboratively working together to store a snapshot | -| region | a subgroup of BOXes within a pool by grouped by geographic proximity to each other | -| data set | a file, directory, or a discrete piece of data sitting in a database | - - -## Pre-conditions -[pre-conditions]: #pre-conditions - -* each BOX is already provisioned with the necessary configuration info in order for the keeper to fully operate -* each keeper in a pool can be trusted to not operate maliciously -* the type of file system that each BOX is backing up is the same - -## Limitations -[limitations]: #limitations - -The following limitations may be encountered while operating a pool: -* file size -* number of files in a directory -* snapshot size -* number of keepers in a pool - -## Configuration - -Configuration data for each node can be split into local and shared. - -### Local - - * local BOX address - * local public/private key - -### Shared - - * remote BOX addresses of participants in the pool - * shared secret - * minimum acceptable replication factor - * warning time (how much time should be given for a warning to be sent out before a limitation occurs) - * has a global default as well as an override for each custom limitation - - -## Local Filesystem -If space permits, the entire contents of a snapshot may be stored on an author's local FS. If the author runs out of space then contents must be sharded across several BOXes. - -If the actual replication factor drops below the minimum acceptable threshold a failure event is dispatched. - -## Conflict Resolution -[conflict-resolution]: #conflict-resolution -A conflict may arise between keepers when a snapshat goes out of sync. This could occur either due to [real time updates](#real-time-updates) or [disk corruption](#disk-corruption). - -### Real-time Updates -[real-time-updates]: #real-time-updates - * more than one person is editing the same data set on several BOXes in a pool at the same time - -### Corruption -[coruption]: #corruption - * a disk becomes corrupt on a BOX - -In either case, the disputing keepers should take the appropriate steps to resolve the conflict. If an appropriate back-out strategy can not be achieved an event should be raised. - -## Events -The following events should be dispatched for an administrative UI. - - * limitation encountered - - * network disruption - - * keeper health - * memory - * disk I/O - * CPU usage - * disk corruption - - * unresolved conflict - - * unacceptable replication factor - -### Event Types - * normal - * warning - * failure - -### Warnings - -Events dispatched before a failure occurs based on a forecasting heuristic to determine how quickly a limit will be reached. - -### Event Frequency -Normal events are dispatched periodically (based on a config param) for historical reporting and warnings|failures are dispatched immediately. - -# Reference-level Explanation -[reference-level-explanation]: #reference-level-explanation - -## Network Architecture - -A peer-peer architecture is used over master/slave so that if a single keeper goes down the rest of the pool will still be able to operate normally. - - * any shared config data is stored on each BOX - - * any shared state required for the retrieval of a data set is store on each BOX - - * no central servers are used for routing - - -## Availability - -@TODO - -Describe some heuristics that can be used to achieve greater availability (ie/ storing entire contents of data set on author) - -### Conflict Resolution - -@TODO - -For detecting / handling conflicts due to data sets going out of sync from real-time updates - -There are different conflict resolution strategies available. - -Both strategies might be used by a single keeper depending on the type of data set or an override. - -### File Integrity Monitoring - -@TODO - -For detecting / handling data set corruption. - -Comparing the contents of chunks on disk with a source of truth. - -The source of truth is only be updated from change events. - - -## Event Dispatcher - -@TODO - -## Data Set Retrieval - -@TODO - -## Regions - -@TODO - -A best effort approach to storing an entire snapshot in a single region so that if a region is severed the pool will still be able to recover the entire snapshot. - -## Network Stack - -@TODO - -Defines how peers will communicate with each other. - -# Drawbacks -[drawbacks]: #drawbacks -Putting the responsibility of data reliability on BOX owners means there is potential for a BOX owner to make a mistake and permanently lose their data. - -# Rationale and alternatives -[rationale-and-alternatives]: #rationale-and-alternatives -An alternative could be to use paid services (cloud storage providers) with their own SLAs to take on the responsibility of data reliability. - -If participating in a decentralized storage network (DSN), the BOX owner could also purchase a mining component to offset their cost. - -There are currently a few drawbacks with this: - - * becoming a storage miner requires a significant upfront investment to cover hardware and staking costs - - * a private pool will always be more efficient since keepers will not have to worry about the overhead of trusting each other - -These options are not mutually exclusive. Offering both options (free and paid) could provide the greatest freedom/flexibility for BOX owners. - -# Prior art -[prior-art]: #prior-art - - * [IPFS cluster](https://cluster.ipfs.io/) - -# Unresolved questions -[unresolved-questions]: #unresolved-questions - -How is a data set reconstructed from chunks? - -How are chunks found on the network? (content discovery) - -What is an ideal replication factor? - -How can contents of entire filesystem be efficiently compared with snapshot (aka source of truth)? - -How can reliability be measured? Markov models? - -How can limits be calculated? - -How can NAT hole punching work in a pnet without any relays? - -Should we consider using a VPN or other alternative such as Tor over libp2p? - -How will bootstrapping work for BOXes not on same LAN? - -Which components can be re-used for data sharing? - -Does data compression need to be taken into account? - -# Future possibilities -[future-possibilities]: #future-possibilities - -Storing a history of the snapshot so an owner can go back in time and recover a data set from a previous state. From da24d0e044718f119cd84ee9cc4aedb2aa12816e Mon Sep 17 00:00:00 2001 From: Aaron Surty Date: Fri, 25 Mar 2022 10:47:10 -0400 Subject: [PATCH 4/5] final first draft personal data reserve (#58) --- docs/RFCs/personal-data-backup.md | 246 ----------------------------- docs/RFCs/personal-data-reserve.md | 246 +++++++++++++++++++++++++++++ 2 files changed, 246 insertions(+), 246 deletions(-) delete mode 100644 docs/RFCs/personal-data-backup.md create mode 100644 docs/RFCs/personal-data-reserve.md diff --git a/docs/RFCs/personal-data-backup.md b/docs/RFCs/personal-data-backup.md deleted file mode 100644 index 73f326b..0000000 --- a/docs/RFCs/personal-data-backup.md +++ /dev/null @@ -1,246 +0,0 @@ -- Feature Name: personal-data-backup -- Start Date: 2022-03-18 -- RFC PR: [functionland/docs/pull/61](https://github.com/functionland/docs/pull/61) -- Functionland Issue: [functionland/docs/issues/58](https://github.com/functionland/docs/issues/58) -- Status: Draft -- Authors: [Aaron Surty](https://github.com/gitaaron), [Farhoud](https://github.com/farhoud) -- Reviewers: @TODO - -# Summary -[summary]: #summary - -This RFC covers how a pool of BOXes can work collaboratively together to improve data reliability. - -# Motivation -[motivation]: #motivation - -A person owns several BOXes and wants a portion of their data replicated across each BOX so that if one of the BOXes malfunctions they should not lose any of their data. - -The following scenarios are handled: - - * adding a new BOX to the pool - - * removing a BOX from the pool - - * adding a hard drive to a BOX - - * removing a hard drive from a BOX - - * a severe network outage occurs severing a region of BOXes from another region - - * data becomes corrupted on a BOX - - * updating the same data set in real time - - -# Guide-level explanation -[guide-level-explanation]: #guide-level-explanation - -## Terminology -[terminology]: #terminology - -| Name | Definition | -|----------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------| -| snapshot | the entire collection of data that might be stored in chunks across several keepers but can be rebuilt to represent the entire file system being backed up | -| keeper | a BOX process responsible for storing a portion of a snapshot and sharing the burden of recreating an entire snapshot | -| BOX | an OS that each keeper runs on | -| author | a BOX where the data set was created or written to last | -| replication factor | how many keepers a chunk of data is stored on; a greater replication factor means greater reliability | -| pool | a group of keepers collaboratively working together to store a snapshot | -| region | a subgroup of BOXes within a pool by grouped by geographic proximity to each other | -| data set | a file, directory, or a discrete piece of data sitting in a database | - - -## Pre-conditions -[pre-conditions]: #pre-conditions - -* each BOX is already provisioned with the necessary configuration info in order for the keeper to fully operate -* each keeper in a pool can be trusted to not operate maliciously -* the type of file system that each BOX is backing up is the same - -## Limitations -[limitations]: #limitations - -The following limitations may be encountered while operating a pool: -* file size -* number of files in a directory -* snapshot size -* number of keepers in a pool - -## Configuration - -Configuration data for each node can be split into local and shared. - -### Local - - * local BOX address - * local public/private key - -### Shared - - * remote BOX addresses of participants in the pool - * shared secret - * minimum acceptable replication factor - * normal event frequency - * warning time - * how much time should be given for a warning to be sent out before a imminent limitation is encountered and a system failure occurs - * has a global default as well as an override for each custom limitation - -## Conflict Resolution -[conflict-resolution]: #conflict-resolution -A conflict may arise between keepers when a snapshat goes out of sync. This could occur either due to [real time updates](#real-time-updates) or [disk corruption](#disk-corruption). - -### Real-time Updates -[real-time-updates]: #real-time-updates - * more than one person is editing the same data set on several BOXes in a pool at the same time - -### Disk Corruption -[disk-corruption]: #disk-corruption - * a disk becomes corrupt on a BOX - -In either case, the disputing keepers will take the appropriate steps to resolve the conflict. If an appropriate back-out strategy can not be achieved, an event is raised. - -## Events -The following events should be dispatched for an administrative UI. - - * limitation imminent - - * limitation encountered - - * keeper health - * memory - * disk I/O - * CPU usage - * disk corruption - - * unresolved conflict - - * unacceptable replication factor - - * keeper added / removed - - * network disruption - - -### Event Types - * normal - * warning - * failure - -### Warnings - -Events dispatched before a failure occurs based on a forecasting heuristic to determine how quickly a limit will be reached. - -### Event Frequency -Normal events are dispatched periodically (based on a config param) for historical reporting and warnings|failures are dispatched immediately. - -## Regions - -If multiple regions are set up, each region contains the entire contents of a snapshot. If a region is severed from the pool, it will still be able to recover the entire snapshot. - -# Reference-level Explanation -[reference-level-explanation]: #reference-level-explanation - -## Network Architecture - -A peer-peer architecture is used over master/slave so that if a single keeper goes down the rest of the pool will still be able to operate normally. - - * any shared config data is stored on each BOX - - * any shared state required for the retrieval of a data set is stored on each BOX - - * no central servers are used for routing - - -## Heuristics - -Some heuristics can be used to achieve greater availability / load times and minimize bandwidth. - -### Local File System First - -If space permits, the entire contents of a snapshot may be stored on an author's local file system. If the author runs out of space then contents must be sharded across several BOXes. - -@TODO - fill out other heuristics - - -### Conflict Resolution - -For detecting / handling conflicts due to data sets going out of sync from real-time updates - -There are different conflict resolution strategies available. - -Both strategies might be used by a single keeper depending on the type of data set or an override. - -### File Integrity Monitoring - -For detecting / handling data set corruption. - -Comparing the contents of chunks on disk with a source of truth. - -The source of truth is only updated from change events. - -## Event Dispatcher - -@TODO - -## Data Set Retrieval - -@TODO - - -## Network Stack - -@TODO - -# Drawbacks -[drawbacks]: #drawbacks -Putting the responsibility of data reliability on BOX owners means there is potential for a BOX owner to make a mistake and permanently lose their data. - -# Rationale and alternatives -[rationale-and-alternatives]: #rationale-and-alternatives -An alternative could be to use paid services (cloud storage providers) with their own SLAs to take on the responsibility of data reliability. - -If participating in a decentralized storage network (DSN), the BOX owner could also purchase a mining component to offset their cost. - -There are currently a few drawbacks with this: - - * becoming a storage miner requires a significant upfront investment to cover hardware and staking costs - - * a private pool will always be more efficient since keepers will not have to worry about the overhead of trusting each other - -These options are not mutually exclusive. Offering both options (free and paid) could provide the greatest freedom/flexibility for BOX owners. - -# Prior art -[prior-art]: #prior-art - - * [IPFS cluster](https://cluster.ipfs.io/) - -# Unresolved questions -[unresolved-questions]: #unresolved-questions - -How is a data set reconstructed from chunks? - -How are chunks found on the network? (content discovery) - -What is an ideal replication factor? - -How can contents of an entire filesystem be efficiently compared with snapshot (aka source of truth)? - -How can reliability be measured? Markov models? - -How can system limits be calculated? - -How can NAT hole punching work in a pnet without any relays? - -Should we consider using a VPN or other alternatives such as Tor over libp2p? - -How will bootstrapping work for BOXes not on same LAN? - -Which components can be re-used for data sharing? - -Does data compression need to be taken into account? - -# Future possibilities -[future-possibilities]: #future-possibilities - -Storing a history of the snapshot so an owner can go back in time and recover a data set from a previous state. diff --git a/docs/RFCs/personal-data-reserve.md b/docs/RFCs/personal-data-reserve.md new file mode 100644 index 0000000..a0af040 --- /dev/null +++ b/docs/RFCs/personal-data-reserve.md @@ -0,0 +1,246 @@ +# Personal Data Reserve +- Start Date: 2022-03-18 +- RFC PR: [functionland/docs/pull/61](https://github.com/functionland/docs/pull/61) +- Functionland Issue: [functionland/docs/issues/58](https://github.com/functionland/docs/issues/58) +- Status: Draft +- Authors: [Aaron Surty](https://github.com/gitaaron), +- Reviewers: [Farhoud](https://github.com/farhoud), [Ehsan](https://github.com/ehsan6sha), [Masih](https://github.com/orgs/functionland/people/masih) + +## Summary +[summary]: #summary + +This RFC covers how a BOX customer's data can be replicated across a group of BOXes physically separated from each other to help prevent unwanted loss. + +## Use Case +[use-case]: #use-case + +A person owns 3-5 BOXes (eg/ one in their home, one at their office and one at a friend's house). + +Their basement floods bricking the one located in their home. + +They are still able to access all of the data that was uploaded to their bricked home BOX from one of their other BOXes or a brand new BOX provisioned with factory settings. + +The following scenarios are handled: + + * adding a new BOX to the reserve + + * removing a BOX from the reserve + + * data becomes corrupted on a BOX + + +## Terminology +[terminology]: #terminology + +| Name | Definition | +|----------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------| +| reserve | a group of keepers collaboratively working together to back up a snapshot | +| snapshot | a collection of data that represents the entire file system being backed up at a given moment in time | +| chunk | a portion of a data set | +| keeper | a BOX process responsible for backing up a file system to the reserve | +| retriever | a BOX process given the task of rebuilding a usable file system from a snapshot | +| author | the BOX where the backed up data set was created | +| peer | a member of the reserve | +| replication factor | how many keepers a chunk of data is stored on; a greater replication factor means greater resilience | +| data set | a file, directory, or data from a database | +| warning time window | how much time should be given to send an alert before an imminent limitation is encountered and a system failure occurs | + + + +## Pre-conditions +[pre-conditions]: #pre-conditions + +* the owner has already authenticated themselves with each BOX + +* the fula-api is persisting data to a disk accessible by the keeper + +* each BOX is already provisioned with the necessary configuration info required to operate a reserve + +## Assumptions + +* each keeper in a reserve can be trusted to not operate maliciously by performing unwanted delete operations + +* the data being backed up is: + * multimedia files such as photos and video + * [orbit db](https://orbitdb.org/) metadata + * plain text files for shared configuration data + +## Out of Scope + +* high-frequency updates to data sets which might be encountered with multi-tenancy or recording streaming video + +* syncing of data between BOXes so that it is usable on each BOX (eg/ CRUD operations on them in real time) + +* conflicts arising from concurrent updates initiated by a user on several different BOXes + +* heuristics used to minimize bandwidth and improve availability + +## Limitations +[limitations]: #limitations + +The following limitations may be encountered while operating a reserve: +* file size +* number of files in a directory +* snapshot size +* number of keepers in a reserve + +## Configuration +[configuration]: #configuration + +Configuration data for each peer can be split into local and shared. + +### Local Configuration + + * local BOX address + + * local public/private key + +### Shared Configuration + + * remote BOX addresses of other peers in the reserve + + * shared secret + + * minimum acceptable replication factor + + * 'normal' event dispatching frequency + + * warning time window + * has a global default as well as an override for each limitation + +## Conflict Resolution +[conflict-resolution]: #conflict-resolution + +Although it can be assumed the snapshot is already free from conflicts caused by concurrent updates initiated from a user, conflicts may still arise from disk corruption or other unforseen keeper errors. + +The disputing keepers will take the appropriate steps to resolve the conflict. If an appropriate resolution can not be achieved, an event is raised. + +## Events +The following events are dispatched for an administrative UI. + + * limitation imminent + + * limitation encountered + + * keeper health + * memory + * disk I/O + * CPU usage + * disk corruption + + * unresolved conflict + + * unacceptable replication factor + + * keeper added / removed + + * network disruption + +### Event Types + * normal + * warning + * failure + +### Warnings + +Events dispatched before a failure occurs based on a forecasting heuristic to determine how quickly a limit will be reached. + +### Event Frequency + +Normal events are dispatched periodically (based on a config param) for historical reporting and warnings|failures are dispatched immediately. + +## Implementation +[reference-level-explanation]: #reference-level-explanation + +### Components + +The components that are needed for this use case are: + + * data set keeper + + * data set retriever + + * event dispatcher + + * file integrity monitor + +### Event Dispatcher + +A mechanism for dispatching events and queueing them to prevent loss if a peer goes down. + +### File Integrity Monitoring + +For detecting / handling disk corruption. + +A full sweep of the file system is periodically done comparing the contents of chunks on disk with a source of truth. + + +### Network Architecture + +A peer-peer architecture is used over master/slave so that if a single BOX goes down the rest of the reserve will still be able to operate normally. + + * any shared config data is stored on each BOX + + * any shared state required for the retrieval of a data set is stored on each BOX + + * no central servers are used for networking + + +## Drawbacks +[drawbacks]: #drawbacks +Putting the responsibility of data reliability on BOX owners means there is potential for a BOX owner to make a mistake and permanently lose their data. + +## Rationale and alternatives +[rationale-and-alternatives]: #rationale-and-alternatives +An alternative could be to use paid services (cloud storage providers) with their own SLAs to take on the responsibility of data reliability. + +If participating in a decentralized storage network (DSN), the BOX owner could also purchase a mining component to offset their cost. + +There are currently a few drawbacks with this: + + * becoming a storage miner requires a significant upfront investment to cover hardware and staking costs + + * a private reserve will always be more efficient since keepers will not have to worry about the overhead of trusting each other + +These options are not mutually exclusive. Offering both options (free and paid) could provide the greatest freedom/flexibility for BOX owners. + +## Dependencies + + * [IPFS cluster](https://cluster.ipfs.io/) + + * [private connections](./pnet) + + * [authentication](./auth) + + * [retriever](./retriever) + + * [keeper](./keeper) + +@TODO - verify links + +## Unresolved questions +[unresolved-questions]: #unresolved-questions + +What is an ideal replication factor? + +How can resilience be measured? Markov models? + +How can contents of an entire filesystem be efficiently compared with a source of truth? + +How can system resources and their limits be calculated? + +How can NAT hole punching work in a pnet without any relays? + +Should data compression be taken into account? + +## Future Possibilities +[future-possibilities]: #future-possibilities + +### History + +Storing a history snapshots so an owner can go back in time and recover a data set from a previous state. This could help with data loss due to user errors. + +### Regions + +A person or group of people may have a large number of BOXes and want to set up geographically grouped regions such that if there is a severe network outage, each region will still be able to recover the entire contents of a snapshot. + From 1f1416e27c98c7ba6cbf832428cf2ade98acfb56 Mon Sep 17 00:00:00 2001 From: Aaron Surty Date: Fri, 25 Mar 2022 12:55:30 -0400 Subject: [PATCH 5/5] fix docusaurus build removing links to docs that do not exist (#58) --- docs/RFCs/personal-data-reserve.md | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/docs/RFCs/personal-data-reserve.md b/docs/RFCs/personal-data-reserve.md index a0af040..5c1ece4 100644 --- a/docs/RFCs/personal-data-reserve.md +++ b/docs/RFCs/personal-data-reserve.md @@ -208,15 +208,15 @@ These options are not mutually exclusive. Offering both options (free and paid) * [IPFS cluster](https://cluster.ipfs.io/) - * [private connections](./pnet) + * private connections - * [authentication](./auth) + * authentication - * [retriever](./retriever) + * retriever - * [keeper](./keeper) + * keeper -@TODO - verify links +@TODO - provide links to relevant docs ## Unresolved questions [unresolved-questions]: #unresolved-questions