ipfs-cluster

Author	SHA1	Message	Date
Hector Sanjuan	10d6a37304	Update deps: fix the things that need fixing License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2019-02-15 12:40:53 +00:00
Adrian Lanzafame	3b3f786d68	add opencensus tracing and metrics This commit adds support for OpenCensus tracing and metrics collection. This required support for context.Context propogation throughout the cluster codebase, and in particular, the ipfscluster component interfaces. The tracing propogates across RPC and HTTP boundaries. The current default tracing backend is Jaeger. The metrics currently exports the metrics exposed by the opencensus http plugin as well as the pprof metrics to a prometheus endpoint for scraping. The current default metrics backend is Prometheus. Metrics are currently exposed by default due to low overhead, can be turned off if desired, whereas tracing is off by default as it has a much higher performance overhead, though the extent of the performance hit can be adjusted with smaller sampling rates. License: MIT Signed-off-by: Adrian Lanzafame <adrianlanzafame92@gmail.com>	2019-02-04 18:53:21 +10:00
Hector Sanjuan	7de930b796	Feat #445 : Use TrackerStatus as filter. Simplify and small misc. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2019-01-07 19:45:07 +01:00
Hector Sanjuan	862c1eb3ea	Fix #382 : Extract headers from IPFS API requests & apply them to hijacked ones. This commit makes the proxy extract useful fixed headers (like CORS) from the IPFS daemon API responses and then apply them to the responses from hijacked endpoints like /add or /repo/stat. It does this by caching a list of headers from the first IPFS API response which has them. If we have not performed any proxied request or managed to obtain the headers we're interested in, this will try triggering a request to "/api/v0/version" to obtain them first. This should fix the issues with using Cluster proxy with IPFS Companion and Chrome. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-12-18 16:05:12 +01:00
Kishan Sagathiya	411eea664c	Merge branch 'master' of github.com:ipfs/ipfs-cluster into issue_453	2018-11-03 20:24:15 +05:30
Kishan Sagathiya	3a5ad6111a	Issue #453 Extract the IPFS Proxy from ipfshttp Changes as requieed Rename IPFSProxy struct to Server License: MIT Signed-off-by: Kishan Mohanbhai Sagathiya <kishansagathiya@gmail.com>	2018-11-01 15:54:05 +05:30
Hector Sanjuan	1bc7f5a643	Re-order shutdownLock fields in cluster declaration License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-10-30 11:54:32 +01:00
Hector Sanjuan	ca3fe646b1	Fix race condition when shutting down and watchPeers() run License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-10-29 12:03:57 +01:00
Hector Sanjuan	765987ef2c	Improvements License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-10-29 12:03:32 +01:00
Hector Sanjuan	85a7dc5114	Allocate memory for response rpc types Might trigger a very weird race when pointing to nil License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-10-29 12:03:03 +01:00
Hector Sanjuan	19b1124999	Make metrics human Issue #572 exposes metrics but they carry the peer ID in binary. This was ok with our internal codecs but it doesn't seem to work very well with json, and makes the output format unusable. This makes the Metric.Peer field a string. Additinoally, fixes calling the command without arguments and displaying the date in the right format. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-10-26 14:11:30 +02:00
Hector Sanjuan	7d16108751	Start using libp2p/go-libp2p-gorpc License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-10-17 15:28:03 +02:00
Kishan Sagathiya	fdb573c96f	Issue #453 Extract the IPFS Proxy from ipfshttp Adding more missing pieces in config Use the right package(not the inbuilt one) Setup rpc client for proxy in the cluster Add back SetClient and Shutdown into Connector as they are required to implement Component interface Add `ipfsproxy` as into list of logging identifier and add its default log level License: MIT Signed-off-by: Kishan Mohanbhai Sagathiya <kishansagathiya@gmail.com>	2018-10-14 22:42:50 +05:30
Kishan Sagathiya	d338f8de53	Issue #453 Extract the IPFS Proxy from ipfshttp Added config Change functions accordingly to add new `ipfsproxy` component License: MIT Signed-off-by: Kishan Mohanbhai Sagathiya <kishansagathiya@gmail.com>	2018-10-14 16:07:42 +05:30
Hector Sanjuan	398cdaa957	Merge pull request #549 from ipfs/fix/543-add-unhealthy Fix #543: Use only healthy peers when adding everywhere (+tests)	2018-09-27 16:52:12 +02:00
Hector Sanjuan	cbdd075bb2	Merge pull request #551 from ipfs/fix/minor-version-compat Allow running peers with different cluster versions in the same cluster	2018-09-27 08:13:53 +02:00
Hector Sanjuan	e7b1eacf83	Fix #543 : Use only healthy peers when adding everywhere (+tests) License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-09-27 08:12:29 +02:00
Hector Sanjuan	59714f69d4	Allow running peers with different cluster versions in the same cluster This patch modifies the RPC protocol tag to use Major and Minor parts of the version and not all of it. This means all peers on the 0.5.x can run in the same cluster. As cluster has become more mature and I see less risks in letting peers from similar versions run together. This is useful when upgrading too. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-09-27 07:35:57 +02:00
Kishan Sagathiya	c62c922e87	Issue #446 Adding peername to PinInfo Avoid typecasting License: MIT Signed-off-by: Kishan Mohanbhai Sagathiya <kishansagathiya@gmail.com>	2018-09-26 08:35:56 +05:30
Kishan Sagathiya	2cd4420ee8	Issue #446 Adding peername to PinInfo Fixing errors and better code placement License: MIT Signed-off-by: Kishan Mohanbhai Sagathiya <kishansagathiya@gmail.com>	2018-09-24 23:20:37 +05:30
Kishan Sagathiya	773b4de1f0	Issue #446 Adding peername to PinInfo Removed comments and code used for debugging License: MIT Signed-off-by: Kishan Mohanbhai Sagathiya <kishansagathiya@gmail.com>	2018-09-24 23:04:16 +05:30
Kishan Sagathiya	ef85ba8780	Issue #446 Adding peername to PinInfo This commit adds peername to PinInfo and GlobalPinInfo so that we have a nicer and more meaningfull output for `ipfs-cluster-ctl` queries like `status`, `sync` and `recover` License: MIT Signed-off-by: Kishan Mohanbhai Sagathiya <kishansagathiya@gmail.com>	2018-09-24 23:04:16 +05:30
Adrian Lanzafame	31474f6490	update go-cid and go-libp2p License: MIT Signed-off-by: Adrian Lanzafame <adrianlanzafame92@gmail.com>	2018-09-24 11:35:38 +10:00
Hector Sanjuan	5bbc699bb4	Issue #340 : Fix some data races Unfortunately, there are still some data races in yamux https://github.com/libp2p/go-libp2p/issues/396 so we can't enable this by default. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-15 12:27:01 +02:00
Hector Sanjuan	1b8967aa46	Merge branch 'master' into feat/sharding-v1 License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-14 11:40:03 +02:00
Hector Sanjuan	f4455d3733	Address comments License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-13 11:56:16 +02:00
Hector Sanjuan	50fc3c4e95	Address comments from review License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-09 13:22:47 +02:00
Hector Sanjuan	0151f5e312	Fixes for adding: set default timeouts to 0. Improve flags and param names. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-08 21:11:26 +02:00
Hector Sanjuan	87f4fcf958	Make the adder modules ipld.DAGService modules. This removes a bunch of the channel dance and block forwarding by having the adder submodules be DAGServices themselves and take Add() directly from the ipfsAdder. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Hector Sanjuan	6ffc8b8347	Update metrics regularly after a few blockputs License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Hector Sanjuan	3575f05753	Support adding Output (api.AddedOutput) This was a long FIXME/TODO. Handling adding output and reporting to the client of the progress of the adding process. This attempts to do it. It is not sure that it works correctly (response body being written while the multipart request is still being read) License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Hector Sanjuan	c81f61eeea	Make sure sync and recover operations receive all cids in a clusterDAG. Cleanup some code and fixmes. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Hector Sanjuan	8f1a15b279	Move adder.Params to api.AddParams. Re-use in other modules License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Hector Sanjuan	aa5589d0d8	codeclimate License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Hector Sanjuan	327a81b85a	golint govets License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Hector Sanjuan	623120fd50	Start cluster tests License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Hector Sanjuan	65dc17a78b	testfixing License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Hector Sanjuan	a96241941e	WIP License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-08-07 20:12:05 +02:00
Wyatt Daviau	82facd3629	fix for pinning same cid sharded and unsharded error License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-08-07 20:11:24 +02:00
Wyatt Daviau	2bc295ca9f	address comments: rebase, cleanup and separate tests, AddParams type replaces map of strings License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-08-07 20:11:24 +02:00
Wyatt Daviau	f7c3dcce5b	removing pointer across rpc instead add logic is a submodule imported in api and cluster comp License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-08-07 20:11:24 +02:00
Wyatt Daviau	1704295331	Refactor add & general cleanup addFile function is now a Cluster method accessed by RPC residue from attempting to stream responses removed ipfs-cluster-ctl ls bug fixed problem with importer/add not printing resolved new test now checks for this License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-08-07 20:11:23 +02:00
Wyatt Daviau	238f3726f3	Pin datastructure updated to support sharding 4 PinTypes specify how CID is pinned Changes to Pin and Unpin to handle different PinTypes Tests for different PinTypes Migration for new state format using new Pin datastructures Visibility of the PinTypes used internally limited by default License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-08-07 20:11:23 +02:00
Wyatt Daviau	e78ccbf6f4	Begin sharding component work: Write basic scaffolding to include a sharding component in cluster Sketch out a high level implementation in pseudo code Share thoughts on upcoming design challenges License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-08-07 20:11:23 +02:00
Hector Sanjuan	d3d1f960f5	Feat: Enable DHT-based peer discovery and routing for cluster peers This uses go-libp2p-kad-dht as routing provider for the Cluster Peers. This means that: * A cluster peer can discover other Cluster peers even if they are not in their peerstore file. * We remove a bunch of code sending and receiving peers multiaddresses when a new peer was added to the Cluster. * PeerAdd now takes an ID and not a multiaddress. We do not need to ask the new peer which is our external multiaddress nor broadcast the new multiaddress to everyone. This will fix problems when bootstrapping a new peer to the Cluster while not all the other peers are online. * Adding a new peer does not mean to open connections to all peers anymore. The number of connections will be made according to the DHT parameters (this is good to have for future work) The that detecting a peer addition in the watchPeers() function does no longer mean that we have connected to it or that we know its multiaddresses. Therefore it's no point to save the peerstore in these events anymore. Here a question opens, should we save the peerstore at all, and should we save multiaddresses only for cluster peers, or for everyone known? Currently, the peerstore is only updated on clean shutdown, and it is updated with all the multiaddresses known, and not limited to peer IDs in the cluster, (because, why not). License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-07-24 15:33:41 +02:00
Adrian Lanzafame	c89508035a	Maptracker: extract optracker and make improvements License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-05-28 11:59:26 +02:00
Hector Sanjuan	4d8f975d9b	StateSync(): some improvements This commit: * Does not collect and return changed items when doing StateSync (they are not used) * Removes the StateSync RPC method (no longer used) * Uses tracker.StatusAll() rather than requesting Status on each Cid (should be faster with upcoming pintracker) * Does not launch a go-routine to track every item. Track is an async operation. This likely causes 1000s goroutines to be started with no good reason. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-05-25 09:58:18 +02:00
Hector Sanjuan	e4844ca819	Monitor: address comments License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-05-09 11:01:52 +02:00
Hector Sanjuan	a9d6fe3479	Types: rename metric.SetTTLDuration to metric.SetTTL GetTTL returns duration. SetTTL should take duration too, not seconds. This removes the original SetTTL method which used seconds. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-05-07 14:26:06 +02:00
Hector Sanjuan	72e1d64de2	Fix publish cancelling contexts too early. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-05-07 14:26:06 +02:00
Hector Sanjuan	3c3341e491	Monitor: add PublishMetric() to component interface The monitor component should be in charge of deciding how it is best to send metrics to other peers and what that means. This adds the PublishMetric() method to the component interface and moves that functionality from Cluster main component to the basic monitor. There is a behaviour change. Before, the metrics where sent only to the leader, while the leader was the only peer to broadcast them everywhere. Now, all peers broadcast all metrics everywhere. This is mostly because we should not rely on the consensus layer providing a Leader(), so we are taking the chance to remove this dependency. Note that in any-case, pubsub monitoring should replace the existing basic monitor. This is just paving the ground. Additionally, in order to not duplicate the multiRPC code in the monitor, I have moved that functionality to go-libp2p-gorpc and added an rpcutil library to cluster which includes useful methods to perform multiRPC requests (some of them existed in util.go, others are new and help handling multiple contexts etc). License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-05-07 14:26:06 +02:00
Hector Sanjuan	33d9cdd3c4	Feat: emancipate Consensus from the Cluster component This commit promotes the Consensus component (and Raft) to become a fully independent thing like other components, passed to NewCluster during initialization. Cluster (main component) no longer creates the consensus layer internally. This has triggered a number of breaking changes that I will explain below. Motivation: Future work will require the possibility of running Cluster with a consensus layer that is not Raft. The "consensus" layer is in charge of maintaining two things: * The current cluster peerset, as required by the implementation * The current cluster pinset (shared state) While the pinset maintenance has always been in the consensus layer, the peerset maintenance was handled by the main component (starting by the "peers" key in the configuration) AND the Raft component (internally) and this generated lots of confusion: if the user edited the peers in the configuration they would be greeted with an error. The bootstrap process (adding a peer to an existing cluster) and configuration key also complicated many things, since the main component did it, but only when the consensus was initialized and in single peer mode. In all this we also mixed the peerstore (list of peer addresses in the libp2p host) with the peerset, when they need not to be linked. By initializing the consensus layer before calling NewCluster, all the difficulties in maintaining the current implementation in the same way have come to light. Thus, the following changes have been introduced: * Remove "peers" and "bootstrap" keys from the configuration: we no longer edit or save the configuration files. This was a very bad practice, requiring write permissions by the process to the file containing the private key and additionally made things like Puppet deployments of cluster difficult as configuration would mutate from its initial version. Needless to say all the maintenance associated to making sure peers and bootstrap had correct values when peers are bootstrapped or removed. A loud and detailed error message has been added when staring cluster with an old config, along with instructions on how to move forward. * Introduce a PeerstoreFile ("peerstore") which stores peer addresses: in ipfs, the peerstore is not persisted because it can be re-built from the network bootstrappers and the DHT. Cluster should probably also allow discoverability of peers addresses (when not bootstrapping, as in that case we have it), but in the meantime, we will read and persist the peerstore addresses for cluster peers in this file, different from the configuration. Note that dns multiaddresses are now fully supported and no IPs are saved when we have DNS multiaddresses for a peer. * The former "peer_manager" code is now a pstoremgr module, providing utilities to parse, add, list and generally maintain the libp2p host peerstore, including operations on the PeerstoreFile. This "pstoremgr" can now also be extended to perform address autodiscovery and other things indepedently from Cluster. * Create and initialize Raft outside of the main Cluster component: since we can now launch Raft independently from Cluster, we have more degrees of freedom. A new "staging" option when creating the object allows a raft peer to be launched in Staging mode, waiting to be added to a running consensus, and thus, not electing itself as leader or doing anything like we were doing before. This additionally allows us to track when the peer has become a Voter, which only happens when it's caught up with the state, something that was wonky previously. * The raft configuration now includes an InitPeerset key, which allows to provide a peerset for new peers and which is ignored when staging==true. The whole Raft initialization code is way cleaner and stronger now. * Cluster peer bootsrapping is now an ipfs-cluster-service feature. The --bootstrap flag works as before (additionally allowing comma-separated-list of entries). What bootstrap does, is to initialize Raft with staging == true, and then call Join in the main cluster component. Only when the Raft peer transitions to Voter, consensus becomes ready, and cluster becomes Ready. This is cleaner, works better and is less complex than before (supporting both flags and config values). We also backup and clean the state whenever we are boostrapping, automatically * ipfs-cluster-service no longer runs the daemon. Starting cluster needs now "ipfs-cluster-service daemon". The daemon specific flags (bootstrap, alloc) are now flags for the daemon subcommand. Here we mimic ipfs ("ipfs" does not start the daemon but print help) and pave the path for merging both service and ctl in the future. While this brings some breaking changes, it significantly reduces the complexity of the configuration, the code and most importantly, the documentation. It should be easier now to explain the user what is the right way to launch a cluster peer, and more difficult to make mistakes. As a side effect, the PR also: * Fixes #381 - peers with dynamic addresses * Fixes #371 - peers should be Raft configuration option * Fixes #378 - waitForUpdates may return before state fully synced * Fixes #235 - config option shadowing (no cfg saves, no need to shadow) License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-05-07 07:39:41 +02:00
Sina Mahmoodi	d8f7a2adcc	cluster: add version diff log to start errors License: MIT Signed-off-by: Sina Mahmoodi <itz.s1na@gmail.com>	2018-04-24 14:27:15 +02:00
Sina Mahmoodi	03cc809708	config: Add log and testcase for disable_repinning * Test case creates a bunch of clusters, assigns a pin with replica factor of n-1 to them, and removes one of the peers randomly. It then tests to check that the number of clusters pinning the cid is n-2. * Add warn log to let user know that due to disable_repinning option, the cluster won't attempt to re-assign the pin. License: MIT Signed-off-by: Sina Mahmoodi <itz.s1na@gmail.com>	2018-04-23 22:01:52 +02:00
Sina Mahmoodi	0954c6d6fa	Add disable_repinning cluster option License: MIT Signed-off-by: Sina Mahmoodi <itz.s1na@gmail.com>	2018-04-22 18:40:46 +02:00
Hector Sanjuan	dd4128affc	Fix #339 : Reduce Sleeps in tests License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-04-05 16:49:26 +02:00
Hector Sanjuan	58acf16efa	cluster: introduce PeerWatchInterval config option. It should provide a way to speed up peer list updates when peers join/part. It was hardcoded. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-04-05 16:49:26 +02:00
Hector Sanjuan	a4adce6592	Merge pull request #349 from ipfs/feat/restapi-libp2p Feat #305: Libp2p support for REST API	2018-03-26 14:22:39 +02:00
Hector Sanjuan	6777122abf	rest/libp2p-http: address @zenground0 comments License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-03-20 19:51:57 +01:00
Hector Sanjuan	09f4c9fce3	rest/libp2p-http: address lanzafame's review License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-03-20 19:35:42 +01:00
Hector Sanjuan	a6acb72a2c	Improve errors for bootstraps License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-03-20 14:27:49 +01:00
Hector Sanjuan	a73d7e6f7e	Relocate multiaddrJoin and multiaddrSplit to api/types.h So they can serve as multi-module helpers without having circular deps. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-03-16 13:37:32 +01:00
Hector Sanjuan	de07a9dd70	Merge pull request #344 from ipfs/feat/wrong-secrets Fix #167: Useful messages when consensus doesn't start	2018-03-16 11:41:05 +01:00
Hector Sanjuan	5956dce69f	Create cluster Host in ipfs-cluster-service. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-03-15 00:04:54 +01:00
Hector Sanjuan	41b17bf477	Cluster: add libp2p host parameter to constructor. NewCluster() now takes an optional Host parameter. The rationale is to allow to re-use an existing libp2p Host when creating the cluster. The NewClusterHost method now allows to create a host with the options used by cluster. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-03-15 00:04:54 +01:00
Hector Sanjuan	4f9fccde72	Addressing feedback License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-03-13 10:37:47 +01:00
Hector Sanjuan	740f314976	Fix #167 : Useful messages when consensus doesn't start This will display a few hints when consensus fails to start. If consensus doesn't start (normally WaitForLeader times out), it's because of libp2p not being able to reach other peers. This sometimes also means that the wrong protector key (secret) is being used, even though libp2p does not give us clear indications. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-03-12 22:54:59 +01:00
Wyatt Daviau	b0e8452020	fix errors License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-03-12 11:33:46 -04:00
Wyatt Daviau	e2c4b6f5a9	consolidate Pin and PinTo License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-03-09 17:16:20 -05:00
Wyatt Daviau	0a34f3382b	ToPin call and priority pinning License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-03-09 17:13:35 -05:00
Hector Sanjuan	5ba746a9ca	Fix: official builds panic on start Since the commit variable is not set in these builds :( License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-02-20 17:55:18 +01:00
Hector Sanjuan	ebd167edc0	Merge pull request #294 from ipfs/fix/jenkins Fix jenkins tests	2018-01-26 12:40:09 +01:00
Wyatt Daviau	eafc747305	fix/297 Resolve the lack of snapshot pushes: Snapshot saving state commands (upgrade and import) now save raft config peers as consensus peers in snapshot. Snapshot index 1 -> 2 when saving from a fresh import to force replication when bootstrapping. License: MIT Signed-off-by: Wyatt Daviau <wdaviau@cs.stanford.edu>	2018-01-25 16:47:12 -05:00
Hector Sanjuan	ddb5da18c9	Tests: Bind testing clusters on random port Jenkins likes this very much. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-01-24 20:16:55 +01:00
Hector Sanjuan	4b6ee706e7	Fix #222 : Fix overpinning or underpinning of pins after rejoin The StateSync() function did not take into account that the maptracker may think that some pinned items are local/remote when they should not be. In those cases, it needs to trigger re-tracks. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-01-19 22:24:03 +01:00
Hector Sanjuan	dcfc962f24	Feat #277 : Address review comments License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-01-19 22:24:03 +01:00
Hector Sanjuan	a1ab106fcc	Feat #277 : Improve getCurretPin() call License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-01-19 22:24:03 +01:00
Hector Sanjuan	ae1afe3af8	Feat #277 : Avoid re-pinning log entries when the new pin is the same This ensures that we don't re-pin something which is already correctly pinned with the same allocations. It also ensures that we do re-pin something when the replication factor associated to it changes. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-01-19 22:24:03 +01:00
Hector Sanjuan	b013850f94	Fear #277 : Add test about wanted < 0 License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-01-19 22:24:03 +01:00
Hector Sanjuan	4549282cba	Fix #277 : Introduce maximum and minimum replication factor This PR replaces ReplicationFactor with ReplicationFactorMax and ReplicationFactor min. This allows a CID to be pinned even though the desired replication factor (max) is not reached, and prevents triggering re-pinnings when the replication factor has not crossed the lower threshold (min). License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-01-16 16:36:06 +01:00
Wyatt	fc237b21d4	Feat: Enable Jenkins builds This enables support for testing in jenkins. Several minor adjustments have been performed to improve the probability that the tests pass, but there are still some random problems appearing with libp2p conections not becoming available or stopping working (similar to travis, but perhaps more often). MacOS and Windows builds are broken in worse ways (those issues will need to be addressed in the future). Thanks to @zenground0 and @victorbjelkholm for support! License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2018-01-11 18:11:46 +01:00
Hector Sanjuan	2b6dfa45cd	cluster-service: add version subcommand and change some startup logging The --version flag is default from our cli library so I left that. The version subcommand prints only the version number + the short commit so it's a bit more easy to parse. I have additionally reduced the amount of output on start up by converting some messages to debug. I wish there was a level between INFO and DEBUG though. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2017-12-13 10:25:01 +01:00
Hector Sanjuan	9a246a237d	fix go vet: address go vet warnings License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2017-12-06 15:15:41 +01:00
Hector Sanjuan	c0628e43ff	fix golint: Address a few golint warnings License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2017-12-06 15:15:38 +01:00
Hector Sanjuan	1e87fccf0e	Merge pull request #257 from te0d/feat/add-peer-identifier Added Hostname Property to Configuration	2017-12-04 14:25:49 +01:00
Hector Sanjuan	d6a7caf7a4	Issue #259 : Address CR comments License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2017-12-04 13:59:48 +01:00
Tom O'Donnell (te0d)	7d43be33a4	Added Peername Configuration Test and Renamed to Peername I've modified the peer identifier to be 'peername'. I've also modified the TestLoadJSON to check that it is correctly read from config and set to a default if empty. Also added 'peername' fields to configurations for various tests.	2017-12-01 13:50:13 -05:00
Hector Sanjuan	4922c95589	Support --local parameter for Status[Local] and Sync[Local] operations This allows to call the Rest API's status and sync endpoints with a "?local=true" parameter. This will trigger operations but only on the local peer. Cluster Local and RPC-Local methods have been accordingly, although they are aliases for the PinTracker methods (but otherwise they would not be exposed in external APIs). ipfs-cluster-ctl has been updated to support the new flag. The rationaly behind this feature is that sometimes, a single cluster peer (or the ipfs daemon in it) is misbehaving. The user then wants to Sync, Recover, or see Status for that single peer. This is specially relevant when working with big pinsets in larger clusters, as a Status() call will be considerably more expensive when broadcasted everywhere. Note that the Rest API keeps returning GlobalPinInfo objects even on local=true calls. This ensures that the user always gets the same datatype from an endpoint. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2017-12-01 12:56:26 +01:00
Tom O'Donnell (te0d)	d1ef3d0493	Fixes for Adding Peer Identifier from Code Review I renamed Hostname to simply Name as to not imply relation to DNS. Removed quotes from formatter, used helper function setting config, and added defensive error check.	2017-11-30 15:21:53 -05:00
Hector Sanjuan	e824aea55e	RecoverAll: Implement RecoverAllLocal() which recovers all pins in a peer This adds API, RPC calls to support RecoverAllLocal() (and expose RecoverLocal() on the Rest API too). cluster-ctl is updated accordingly. License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2017-11-30 01:53:31 +01:00
Tom O'Donnell (te0d)	c6c8512a27	Added Hostname Property to Configuration I added a "hostname" property to a node's configuration file. Its value defaults to the hostname provided by the OS, but can be modified to anything beside an empty string in the config file. The "hostname" was added to the output of the "id" call. Thus, peer hostnames are available when listing peers.	2017-11-29 15:44:31 -05:00
Wyatt	47b744f1c0	ipfs-cluster-service state upgrade cli command ipfs-cluster-service now has a migration subcommand that upgrades persistant state snapshots with an out-of-date format version to the newest version of raft state. If all cluster members shutdown with consistent state, upgrade ipfs-cluster, and run the state upgrade command, the new version of cluster will be compatible with persistent storage. ipfs-cluster now validates its persistent state upon loading it and exits with a clear error in the case the state format version is not up to date. Raft snapshotting is enforced on all shutdowns and the json backup is no longer run. This commit makes use of recent changes to libp2p-raft allowing raft states to implement their own marshaling strategies. Now mapstate handles the logic for its (de)serialization. In the interest of supporting various potential upgrade formats the state serialization begins with a varint (right now one byte) describing the version. Some go tests are modified and a go test is added to cover new ipfs-cluster raft snapshot reading functions. Sharness tests are added to cover the state upgrade command.	2017-11-28 22:35:48 -05:00
Hector Sanjuan	1f93662b3e	cluster: get first peerset from configuration make sure we save a new config if the new peerset is different than the one in the configuration at boot. Hopefully this fixes a race condition in PeerAdd test License: MIT Signed-off-by: Hector Sanjuan <code@hector.link>	2017-11-15 18:01:49 +01:00
Hector Sanjuan	a656e45375	cluster: safeguard consensus not set when calling ID SwarmConnect on the ipfs connector calls rpc Peers() which requests IDs for every peer member. If that peer member is booting, it might get the request after RPC is setup but before consensus is initialized. In which case a panic happens. Probability that this happens is small, but still. Also increase the connect swarms delay to 30 seconds, which should be a bit longer than the default wait_for_leader timeout, otherwise we might connect swarms while there's not even a leader. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-15 16:38:21 +01:00
Hector Sanjuan	76fe62fae4	Fix error message when not enough candidates for pinning exist License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-15 03:17:08 +01:00
Hector Sanjuan	145dced3e8	Cluster: Fix libp2p host getting shutdown in the middle of peer removal This is what it was likely causing PeerRemove tests to fail randomly but very often. We cancelled the Cluster context before shutting down the Consensus component. This killed networking and aborted the peer remove operations when the leader is removing itself. As a result, it would error with "leadership lost", which would trigger a retry which would set the final error to "context cancelled" because the shutdown of the consensus component proceeds during the retry, cancelling the consensus context. This is not only affecting tests, it might affected operations when running cluster. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-15 02:33:46 +01:00
Hector Sanjuan	cc81ffe96b	cluster: peerAdd: try to return an up-to-date new peer ID Sometimes tests fail because the returned api.ID for a new peer does not include the current cluster peers. This is because the new peerset has not yet be commited in the new peer at the time of the request. This commit retries obtaining the request until the correct peerset comes in, or gives up after two seconds retrying. Rather than the tests failing, note that the ID returned it is very user-facing and should contain the current cluster peers after adding, and not the former peerset, at least while peerAdd operation is allowed. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-14 23:54:23 +01:00
Hector Sanjuan	5e465c3f62	PeerAdd: send cluster multiaddresses to new peer before adding it to raft It seems more logical that the new peer should know how to contact everyone before it needs to start doing it, but it probably does not matter much. Still, more logical. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-14 21:04:33 +01:00
Hector Sanjuan	01fc55550a	Issue #219 : make sure peers are saved after Join License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-10 16:22:31 +01:00
Hector Sanjuan	2a616aeddb	Issue #219 : Provide a list of peers and a list of addresses in the ID object. Fix cluster-ctl to show right number of peers License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-10 16:11:09 +01:00
Hector Sanjuan	b852dfa892	Fix #219 : WIP: Remove duplicate peer accounting This change removes the duplicities of the PeerManager component: * No more commiting PeerAdd and PeerRm log entries * The Raft peer set is the source of truth * Basic broadcasting is used to communicate peer multiaddresses in the cluster * A peer can only be added in a healthy cluster * A peer can be removed from any cluster which can still commit * This also adds support for multiple multiaddresses per peer License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-08 20:04:04 +01:00
Hector Sanjuan	bff1ec3635	Issue #131 : rename addFromMultiaddrs to setFromMultiaddrs License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-01 19:38:46 +01:00
Hector Sanjuan	073c43e291	Issue #131 : Make sure peers are moved to bootstrap when leaving Also, do not shutdown when seeing our own departure during bootstrap. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-01 13:58:57 +01:00
Hector Sanjuan	c912cfd205	Issue #131 : Destroy raft data when the peer has been removed License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-01 13:25:28 +01:00
Hector Sanjuan	7a5f8f184b	Issue #131 : Improvements adding and removing This works on remove+shutdown procedure and fixes a few small issues. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-01 13:00:32 +01:00
Hector Sanjuan	10c7afbd59	Raft: re-enable: do not start with unconsistent peers. Improved error messages License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-01 12:17:33 +01:00
Hector Sanjuan	107ad3bdfc	Shutdown: do not save peers if we did not become ready License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-01 12:17:33 +01:00
Hector Sanjuan	7540e7b056	Leave on shutdown: only attempt when cluster reached ready state. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-01 12:17:33 +01:00
Hector Sanjuan	0980a5de70	Issue #131 : Do not crash when shutting down after consensus start error License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-01 12:17:33 +01:00
Hector Sanjuan	848023e381	Fix #139 : Update cluster to Raft 1.0.0 The main differences is that the new version of Raft is more strict about starting raft peers which already contain configurations. For a start, cluster will fail to start if the configured cluster peers are different from the Raft peers. The user will have to manually cleanup Raft (TODO: an ipfs-cluster-service command for it). Additionally, this commit adds extra options to the consensus/raft configuration section, adds tests and improves existing ones and improves certain code sections. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-11-01 12:17:33 +01:00
Hector Sanjuan	828236dcc0	Issue #213 : Make sure we wait for configuration to be saved There might be a case where the program is terminated before configuration is saved. Also, avoid calling save() multiple times on shutdowns. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-10-27 21:44:02 +02:00
Hector Sanjuan	8f06baa1bf	Issue #162 : Rework configuration format The following commit reimplements ipfs-cluster configuration under the following premises: * Each component is initialized with a configuration object defined by its module * Each component decides how the JSON representation of its configuration looks like * Each component parses and validates its own configuration * Each component exposes its own defaults * Component configurations are make the sections of a central JSON configuration file (which replaces the current JSON format) * Component configurations implement a common interface (config.ComponentConfig) with a set of common operations * The central configuration file is managed by a config.ConfigManager which: * Registers ComponentConfigs * Assigns the correspondent sections from the JSON file to each component and delegates the parsing * Delegates the JSON generation for each section * Can be notified when the configuration is updated and must be saved to disk The new service.json would then look as follows: ```json { "cluster": { "id": "QmTVW8NoRxC5wBhV7WtAYtRn7itipEESfozWN5KmXUQnk2", "private_key": "<...>", "secret": "00224102ae6aaf94f2606abf69a0e278251ecc1d64815b617ff19d6d2841f786", "peers": [], "bootstrap": [], "leave_on_shutdown": false, "listen_multiaddress": "/ip4/0.0.0.0/tcp/9096", "state_sync_interval": "1m0s", "ipfs_sync_interval": "2m10s", "replication_factor": -1, "monitor_ping_interval": "15s" }, "consensus": { "raft": { "heartbeat_timeout": "1s", "election_timeout": "1s", "commit_timeout": "50ms", "max_append_entries": 64, "trailing_logs": 10240, "snapshot_interval": "2m0s", "snapshot_threshold": 8192, "leader_lease_timeout": "500ms" } }, "api": { "restapi": { "listen_multiaddress": "/ip4/127.0.0.1/tcp/9094", "read_timeout": "30s", "read_header_timeout": "5s", "write_timeout": "1m0s", "idle_timeout": "2m0s" } }, "ipfs_connector": { "ipfshttp": { "proxy_listen_multiaddress": "/ip4/127.0.0.1/tcp/9095", "node_multiaddress": "/ip4/127.0.0.1/tcp/5001", "connect_swarms_delay": "7s", "proxy_read_timeout": "10m0s", "proxy_read_header_timeout": "5s", "proxy_write_timeout": "10m0s", "proxy_idle_timeout": "1m0s" } }, "monitor": { "monbasic": { "check_interval": "15s" } }, "informer": { "disk": { "metric_ttl": "30s", "metric_type": "freespace" }, "numpin": { "metric_ttl": "10s" } } } ``` This new format aims to be easily extensible per component. As such, it already surfaces quite a few new options which were hardcoded before. Additionally, since Go API have changed, some redundant methods have been removed and small refactoring has happened to take advantage of the new way. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-10-18 00:00:12 +02:00
Wyatt	e3ccc1b8f4	Using unshadow to save bootstrappers without changing other functionality	2017-10-11 16:12:21 -04:00
Wyatt	67d38a06c4	Using shadow to actually save bootstrapper, updating cluster restart to respect saved config for tests	2017-10-11 11:09:39 -04:00
Wyatt	a1ec459b30	Peers saved in bootstrapper upon peer rm	2017-10-07 20:27:36 +03:00
Hector Sanjuan	8d3c72b766	Fix tests: Make metric broadcasting async When a peer is down, metric broadcasting hangs, and no more ticks are sent for a while.	2017-07-21 23:45:33 +02:00
Hector Sanjuan	ab1cc47d75	Fix #97 : Assume default DataFolder as subfolder to config folder when empty. We no longer set ConsensusDataFolder. We leave it empty (and ommited from the configuration). When not set, it will take the path from which the configuration file was read and use an "ipfs-cluster-data" subfolder in that path. When set, the behaviour is just as before (ensures backwards compatiblity). This will facilitate re-use of configuration files, for example, when mounting them inside docker. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-07-18 11:36:24 +02:00
Hector Sanjuan	fc029269ff	Merge pull request #109 from ipfs/feat/pnet Private Network impl	2017-07-13 21:33:55 +02:00
dgrisham	98335901fc	Refactored private network implementation + config.	2017-07-08 11:11:49 -06:00
Hector Sanjuan	157e98d25e	Print number of candidates when not getting enough License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-07-05 19:08:37 +02:00
Hector Sanjuan	ec0451d9e8	Print current candidates in the error when there are not enough License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-07-05 18:53:05 +02:00
Hector Sanjuan	73b7dc53a1	Fix tests License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-07-05 17:00:19 +02:00
Hector Sanjuan	b1b5bed544	allocate: shortcut when not enough candidates repinFromPeer: use helper function. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-07-05 16:38:36 +02:00
Hector Sanjuan	faa755f43a	Re-allocate pins on peer removal PeerRm now triggers re-pinning of all the Cids allocated to the removed peer. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-07-05 16:38:36 +02:00
dgrisham	59fde30e1e	swarm secret implementation started	2017-07-03 14:00:01 -06:00
dgrisham	1d90130f65	initial tests passing	2017-06-29 18:59:36 -06:00
dgrisham	7bcf64f6b5	build succeeds, PNETs seem to work -- still need tests License: MIT Signed-off-by: David Grisham <dgrisham@mines.edu>	2017-06-27 10:30:15 -06:00
Hector Sanjuan	d0fff1022d	Fix #105 : Panic when calling globalPinInfoCid and peers are down We were setting a nil Cid on the PinInfo object for errors. Cleaned up the logic a bit, and added comments License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-06-21 14:24:23 +02:00
Hector Sanjuan	98039bdd0d	Merge pull request #91 from ipfs/pin-ls-cid Fix #87: Implement ipfs-cluster-ctl pin ls <cid>	2017-04-07 01:35:01 +02:00
Hector Sanjuan	bb82c27b25	Fix #87 : Implement ipfs-cluster-ctl pin ls <cid> I have updated API endpoints to be /allocations rather than /pinlinst It's more self-explanatory. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-04-06 21:12:16 +02:00
Hector Sanjuan	856252e5c6	Fix #88 : Run SyncAllLocal() regularly on peers. It makes a pin ls requests to ipfs and makes sure the pin tracker is up to date. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-04-05 23:29:22 +02:00
Hector Sanjuan	18034356fa	Fix #57 : Hijack /add requests and pin items in cluster after ipfs adds them. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-30 11:27:40 +02:00
Hector Sanjuan	43aa5ffa1f	Fix: do not bootstrap twice :S License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-28 17:27:45 +02:00
Hector Sanjuan	4bb30cd24a	Fixes #16 : trigger ipfs swarm connect to other ipfs nodes in the cluster. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-27 12:42:54 +02:00
Hector Sanjuan	e2efef8469	go lint, go vet, put the Consensus component behind interface. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-14 16:37:29 +01:00
Hector Sanjuan	a40e90a78d	Fix: metrics with nanoseconds TTLs It turns out they only worked in round seconds. Tests send a metric every second, so sometimes they were expired right away. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-14 14:18:23 +01:00
Hector Sanjuan	c2faf48177	Issue #18 : Move Consensus and PeerMonitor to its own submodules License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-13 18:40:35 +01:00
Hector Sanjuan	718b2177ce	Issue #51 : Save a backup on shutdown This adds snapshot and restore methods to state and uses the snapshot one to save a copy of the state when shutting down. Right now, this is not used for anything else. Some lines performing a migration, but this is only an idea of how it could work. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-13 17:57:10 +01:00
Hector Sanjuan	acefb68c8a	Only leader broadcasts metrics. The rest sends only to leader This is a better approach than broadcasting everything all the time (see `6ee0f3bead`). A couple of delays have been touched in order to make tests less likely to fail randomly. (Issue #65). License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-13 15:55:52 +01:00
Hector Sanjuan	a1d31d094b	Fix pin re-allocation when it's already pinned There was a bug in the test for re-allocation which hid another bug in the re-allocation process where a pin which needs to be re-allocated would only be assigned to the new destinations and not anymore to the valid allocations it had previously License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-09 14:49:10 +01:00
Hector Sanjuan	01d65a1595	Support replication factor as a pin parameter This adds a replication_factor query argument to the API endpoint which allows to set a replication factor per Pin. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-08 18:50:54 +01:00
Hector Sanjuan	9b652bcfb3	Rename CidArg to Pin. CidArg used to be an internal name for an argument that carried a Cid. Now it has surfaced to API level and makes no sense. It is a Pin. It represents a Pin (Cid, Allocations, Replication Factor) License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-08 16:57:27 +01:00
Hector Sanjuan	6ee0f3bead	Issue #45 : Detect expired metrics and trigger re-pins An initial, simple approach to this. The PeerMonitor will check it's metrics, compare to the current set of peers and put an alert in the alerts channel if the metrics for a peer have expired. Cluster reads this channel looking for "ping" alerts. The leader is in charge of triggering repins in all the Cids allocated to a given peer. Also, metrics are now broadcasted to the cluster instead of pushed only to the leader. Since they happen every few seconds it should be okay regarding how it scales. Main problem was that if the leader is the node going down, the new leader will not now about it as it doesn't have any metrics for it, so it won't trigger an alert. If it acted on that then the component needs to know it is the leader, or cluster needs to handle alerts in complicated ways when leadership changes. Detecting leadership changes or letting a component know who is the leader is another dependency from the consensus algorithm that should be avoided. Therefore we broadcast, for the moment. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-03-02 14:59:45 +01:00
Hector Sanjuan	37046dc925	go vet fixes License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-02-15 15:46:51 +01:00
Hector Sanjuan	56d68dded1	Avoid showing duplicate Addresses in peers IDs.Addresses. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-02-15 14:15:12 +01:00
Hector Sanjuan	f935eb4245	Use c.id shorthand instead of c.host.ID() License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-02-15 14:15:12 +01:00
Hector Sanjuan	2512ecb701	Issue #41 : Add Replication factor New PeerManager, Allocator, Informer components have been added along with a new "replication_factor" configuration option. First, cluster peers collect and push metrics (Informer) to the Cluster leader regularly. The Informer is an interface that can be implemented in custom wayts to support custom metrics. Second, on a pin operation, using the information from the collected metrics, an Allocator can provide a list of preferences as to where the new pin should be assigned. The Allocator is an interface allowing to provide different allocation strategies. Both Allocator and Informer are Cluster Componenets, and have access to the RPC API. The allocations are kept in the shared state. Cluster peer failure detection is still missing and re-allocation is still missing, although re-pinning something when a node is down/metrics missing does re-allocate the pin somewhere else. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-02-14 19:13:08 +01:00
Hector Sanjuan	0e7091c6cb	Move testing mocks to subpackage so they can be re-used Related to #18 License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-02-09 17:51:19 +01:00
Hector Sanjuan	1b3d04e18b	Move all API-related types to the /api subpackage. At the beginning we opted for native types which were serializable (PinInfo had a CidStr field instead of Cid). Now we provide types in two versions: native and serializable. Go methods use native. The rest of APIs (REST/RPC) use always serializable versions. Methods are provided to convert between the two. The reason for moving these out of the way is to be able to re-use type definitions when parsing API responses in `ipfs-cluster-ctl` or any other clients that come up. API responses are just the serializable version of types in JSON encoding. This also reduces having duplicate types defs and parsing methods everywhere. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-02-09 16:30:53 +01:00
Hector Sanjuan	34fdc329fc	Fix #24 : Auto-join and auto-leave operations for Cluster This is the third implementation attempt. This time, rather than broadcasting PeerAdd/Join requests to the whole cluster, we use the consensus log to broadcast new peers joining. This makes it easier to recover from errors and to know who exactly is member of a cluster and who is not. The consensus is, after all, meant to agree on things, and the list of cluster peers is something everyone has to agree on. Raft itself uses a special log operation to maintain the peer set. The tests are almost unchanged from the previous attempts so it should be the same, except it doesn't seem possible to bootstrap a bunch of nodes at the same time using different bootstrap nodes. It works when using the same. I'm not sure this worked before either, but the code is simpler than recursively contacting peers, and scales better for larger clusters. Nodes have to be careful about joining clusters while keeping the state from a different cluster (disjoint logs). This may cause problems with Raft. License: MIT Signed-off-by: Hector Sanjuan <hector@protocol.ai>	2017-02-07 18:46:09 +01:00

1 2 3 4 5 ...

293 Commits