<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Zookeeper on Brave New Geek</title><link>https://bravenewgeek.com/tag/zookeeper/</link><description>Recent content in Zookeeper on Brave New Geek</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 23 Feb 2018 16:10:02 -0600</lastBuildDate><atom:link href="https://bravenewgeek.com/tag/zookeeper/index.xml" rel="self" type="application/rss+xml"/><item><title>Building a Distributed Log from Scratch, Part 3: Scaling Message Delivery</title><link>https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-3-scaling-message-delivery/</link><pubDate>Mon, 08 Jan 2018 16:10:40 -0600</pubDate><guid>https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-3-scaling-message-delivery/</guid><description>&lt;p&gt;In &lt;a href="https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/"&gt;part two&lt;/a&gt; of this series we discussed data replication within the context of a distributed log and how it relates to high availability. Next, we’ll look at what it takes to scale the log such that it can handle non-trivial workloads.&lt;/p&gt;
&lt;h3 id="data-scalability"&gt;Data Scalability&lt;/h3&gt;
&lt;p&gt;A key part of scaling any kind of data-intensive system is the ability to partition the data. Partitioning is how we can scale a system linearly, that is to say we can handle more load by adding more nodes. We make the system &lt;em&gt;horizontally&lt;/em&gt; scalable.&lt;/p&gt;</description></item><item><title>Building a Distributed Log from Scratch, Part 2: Data Replication</title><link>https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/</link><pubDate>Wed, 27 Dec 2017 12:26:55 -0600</pubDate><guid>https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/</guid><description>&lt;p&gt;In &lt;a href="https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-1-storage-mechanics/"&gt;part one&lt;/a&gt; of this series we introduced the idea of a message log, touched on why it’s useful, and discussed the storage mechanics behind it. In part two, we discuss data replication.&lt;/p&gt;
&lt;p&gt;We have our log. We know how to write data to it and read it back as well as how data is persisted. The caveat to this is, although we have a durable log, it’s a single point of failure (SPOF). If the machine where the log data is stored dies, we’re SOL. Recall that one of our three priorities with this system is high availability, so the question is how do we achieve high availability and fault tolerance?&lt;/p&gt;</description></item><item><title>Benchmarking Commit Logs</title><link>https://bravenewgeek.com/benchmarking-commit-logs/</link><pubDate>Sun, 27 Nov 2016 13:28:55 -0600</pubDate><guid>https://bravenewgeek.com/benchmarking-commit-logs/</guid><description>&lt;p&gt;In this article, we look at &lt;a href="https://kafka.apache.org/"&gt;Apache Kafka&lt;/a&gt; and &lt;a href="http://nats.io/"&gt;NATS Streaming&lt;/a&gt;, two messaging systems based on the idea of a commit log. We’ll compare some of the features of both but spend less time talking about Kafka since by now it’s quite well known. Similar to &lt;a href="https://bravenewgeek.com/benchmarking-message-queue-latency/"&gt;previous&lt;/a&gt; &lt;a href="https://bravenewgeek.com/dissecting-message-queues/"&gt;studies&lt;/a&gt;, we’ll attempt to quantify their general performance characteristics through careful benchmarking.&lt;/p&gt;
&lt;p&gt;The purpose of this benchmark is to test drive the newly released NATS Streaming system, which was made generally available just in the last few months. NATS Streaming doesn’t yet support clustering, so we try to put its performance into context by looking at a similar configuration of Kafka.&lt;/p&gt;</description></item><item><title>You Cannot Have Exactly-Once Delivery</title><link>https://bravenewgeek.com/you-cannot-have-exactly-once-delivery/</link><pubDate>Wed, 25 Mar 2015 18:14:57 -0500</pubDate><guid>https://bravenewgeek.com/you-cannot-have-exactly-once-delivery/</guid><description>&lt;p&gt;I’m often surprised that people continually have fundamental misconceptions about how distributed systems behave. I myself shared many of these misconceptions, so I try not to demean or dismiss but rather educate and enlighten, hopefully while sounding less preachy than that just did. I continue to learn only by following in the footsteps of others. In retrospect, it shouldn’t be surprising that folks buy into these fallacies as I once did, but it can be frustrating when trying to communicate certain design decisions and constraints.&lt;/p&gt;</description></item><item><title>Not Invented Here</title><link>https://bravenewgeek.com/not-invented-here/</link><pubDate>Sat, 06 Dec 2014 14:17:52 -0600</pubDate><guid>https://bravenewgeek.com/not-invented-here/</guid><description>&lt;p&gt;Engineers love engineering things. The reason is self-evident (and maybe self-fulfilling—why else would you be an engineer?). We like to think we’re pretty good at solving problems. Unfortunately, this mindset can, on occasion, yield undesirable consequences which might not be immediately apparent but all the while damaging.&lt;/p&gt;
&lt;p&gt;Developers are all in tune with the idea of “don’t reinvent the wheel,” but it seems to be eschewed sometimes, deliberately or otherwise. People don’t generally write their own merge sort, so why would they write their own consensus protocol? Anecdotally speaking, &lt;a href="https://groups.google.com/d/msg/redis-db/Oazt2k7Lzz4/MPhvXVizCAYJ"&gt;they do&lt;/a&gt;.&lt;/p&gt;</description></item></channel></rss>