
Software Engineering Daily
2,221 episodes — Page 40 of 45

Ep 322Apache Arrow with Uwe Korn
In a typical data analytics system, there are a variety of technologies interacting. HDFS for storing files, Spark for distributed machine learning, pandas for data analysis in Python–each of these different technologies has a different format for how data is represented. Serialization and deserialization between these different formats causes significant latency across the overall system. Apache Arrow is a tool for improving performance of in-memory analytics systems, and today’s guest Uwe Korn explains how Arrow enables these systems with interoperability.

Ep 321Economics of Software with Russ Roberts
EconTalk is a weekly economics podcast that has been going for a decade. On EconTalk, Russ Roberts brings on writers, intellectuals, and entrepreneurs for engaging conversations about the world as seen through the lens of economics. Russ Roberts is today’s guest, and it is a treat because I have been listening to EconTalk since 2006 and it was a central point of inspiration for what Software Engineering Daily has become. On this episode, we talk about how software impacts the world economically, from bitcoin’s promise of zero cost transactions to the opportunities and regulatory challenges of the software-enabled gig economy.

Ep 320IoT Analytics with Jean-Christophe Cimono
On smart thermostats, sensor-driven assembly lines, and electronically monitored farms, the internet of things is producing huge volumes of data. To take advantage of that data, an application needs tools for storing and analyzing that data. Today’s guest is Jean-Cristophe Cimono, the CTO of mnubo, a cloud platform for connected objects. Today we walk through the architecture of mnubo and the use cases of an IoT analytics platform. Thanks to Manuel Vonthron and Eduardo Siman for contributing to the preparation of this show.
Cassandra Data Modeling with Jon Haddad
Salary Negotiation with Haseeb Qureshi

Ep 317Platforms with Bridget Kromhout
At software conferences, I like to walk around the vendor booths and talk to the representatives from different companies. By talking to the vendors about their marketing pitches, I get an idea of how those companies are positioning themselves for the future, and the complex business landscape of software becomes slightly easier to understand. At recent conferences, many of the big vendors have been talking about their cloud platform. With Cloud Foundry, OpenShift, Kubernetes, OpenStack and so many other cloud platforms, it is hard to keep track of the different offerings, and how they differentiate. Bridget Kromhout joins the show today to talk about these different platforms, and how they have changed modern software development. We also talk about her podcast Arrested DevOps, a great show where she interviews some of the luminaries of the operations and software development world. You can also hear Bridget speak at O’Reilly’s Velocity 2016 conference in New York. We are giving away a free ticket to Velocity New York. If you want to be entered to win that ticket, send a tweet about your favorite SE Daily episode about dev ops or web performance, and tag software_daily, as well as #velocityconf.

Ep 316Scalable Architecture with Lee Atchison
Lee Atchison spent seven years at Amazon working in retail, software distribution, and Amazon Web Services. He then moved to New Relic, where he has spent four years scaling the company’s internal architecture. From his decade of experience at fast growing web technology companies, Lee has written the book Architecting for Scale, from O’Reilly. As an application scales, it becomes significantly more complicated while at the same time receiving more traffic. The intersection of these two problems leads to a variety of discussions around availability, risk management, and microservices. Lee and I didn’t have time to get through everything in his book Architecting for Scale, but if you enjoy this episode, check out the book. Lee also spoke recently at the O’Reilly Velocity conference in Santa Clara, so you can check out his talk. We are giving away a free ticket to O’Reilly’s Velocity 2016 conference in New York. If you want to be entered to win that ticket, send a tweet about your favorite SE Daily episode about dev ops or web performance, and tag software_daily, as well as #velocityconf.

Ep 315Schedulers with Adrian Cockcroft
Scheduling is the method by which work is assigned to resources to complete that work. At the operating system level, this can mean scheduling of threads and processes. At the data center level, this can mean scheduling Hadoop jobs or other workflows that require the orchestration of a network of computers. Adrian Cockcroft worked on scheduling at Sun Microsystems, eBay, and Netflix. In each of these environments, the nature of what was being scheduled was different, but the goals of the scheduling algorithms were analogous–throughput, response time, and cache affinity are relevant in different ways at each layer of the stack. Adrian is well-known for helping bring Netflix onto Amazon Web Services, and I recommend watching the numerous YouTube videos of Adrian talking about that transformation.

Ep 314Security and Machine Learning in the Call Center with Pindrop Security’s Chris Halaschek
Call centers are a vulnerable point of attack for large enterprises. Fraud accounts for more than $20 billion in lost money every year, and a significant portion of that fraud is due to customer service representatives being fraudulent social engineering attacks. Chris Halaschek joins the show today to discuss how Pindrop Security is addressing this attack vector. Every phone call that gets made to a call center has a unique phoneprint, and the machine learning model at Pindrop Security uses these phoneprints to assign a risk score to each call. Chris also discusses the challenges associated with scaling a cloud security company.
Cloud Providers with Don Pezet

Ep 312KubeCloud: Tangible Cloud Computing with Kasper Nissen and Martin Jensen
At most universities, there is not a course titled “cloud computing”. Most students leave college without an understanding of distributed systems, cloud service providers, and the fundamentals of how a data center works. Kasper Nissen and Martin Jensen are changing that with KubeCloud, a small tangible cloud computing cluster that runs on Raspberry Pis. Kasper and Martin started KubeCloud as a masters thesis, and it is grown to a textbook-sized treatise on cloud computing. KubeCloud is both software and a curriculum to teach students microservices, containers management, and the real-world problems of distributed systems.
Container Management with Alexis Richardson

Ep 310P2P Money Transfer with TransferWise’s Harsh Sinha
Transferring money from one country to another is expensive, and the banks that facilitate money transfer have tricked us into believing that it should be expensive. On today’s show, Harsh Sinha explains the peer-to-peer system of transferring money with TransferWise, where he works as VP of engineering. Harsh also discusses the larger picture of FinTech companies. The emergence of so many companies at the intersection of finance and technology is no accident. The 2008 financial crisis created a loss of trust in the existing financial system. Simultaneously, smart phones and cheap cloud computing has created opportunities for newer companies like TransferWise to position themselves as a new option for consumer banking.

Ep 309Cassandra Compliant ScyllaDB with Dor Laor
Apache Cassandra is a distributed database that can handle large amounts of data with no single point of failure. Since 2008, Cassandra has been widely adopted and the software and the community around it have grown steadily. A software developer interacting with Cassandra uses CQL, the Cassandra Query Language. ScyllaDB is another open-source database that has been created to be totally compatible with CQL. By complying with CQL, the internals of ScyllaDB can be a vastly different rewrite from Cassandra. ScyllaDB uses C++, whereas Cassandra uses Java. ScyllaDB improves upon the performance characteristics of Cassandra, by optimizing for modern hardware, and Dor Laor joins the show today to discuss how ScyllaDB does all of this.

Ep 308Apache Guacamole and Remote Desktop with Mike Jumper
In order to use a remote desktop experience, software engineers have a limited number of options, and most of them are proprietary, like VMWare or Oracle. Remote desktop is a functionality that many engineers use every day, so it is surprising that the open source world has taken awhile to displace the functionality of proprietary software. In 2010, Mike Jumper started working on Guacamole, a way to access remote desktops through your browser. Over the last six years, Mike has worked continuously to create a simple, open-source software tool to access desktops remotely, and this year Guacamole joined the Apache Incubator and became Apache Guacamole. In this episode, we discuss the past, present, and future of remote desktop, and the technical internals of Apache Guacamole.
Cloud.gov with Aidan Feldman
Death and Distributed Systems with Pieter Hintjens

Ep 305Scaling Twitter with Buoyant.io’s William Morgan
Six years ago, Twitter was experiencing outages due to high traffic. Back in 2010 Twitter was built as a monolithic Ruby on Rails application. Twitter migrated to a microservices architecture to fix these problems. During this migration, the engineers at Twitter learned how to build and scale highly distributed microservice architectures. William Morgan was an engineer at Twitter during that time, and he is now the CEO of Buoyant.io, a company building open-source microservices infrastructure. Some of the big problems at Twitter were solved at the communication layer, using an RPC library called Finagle. At Buoyant, those lessons are being applied to a project called Linkerd, an RPC proxy.

Ep 304Manufacturing and Microservices with Cimpress’ Jim Sokoloff and Maarten Wensveen
Mass customization is the process of making customized, personalized products that are accessible to individuals and small businesses. The process involves manufacturing, assembly lines, supply chains, and software at every step along the way. Today’s guests are Jim Sokoloff and Maarten Wensveen, who work on infrastructure and technology at Cimpress, a mass customization platform. Cimpress has t shirt printers, warehousing machines, supply chain management tools, and lots of other computers that come together in the computer-integrated manufacturing process. The company has been around for a few decades, and more recently they have moved to microservices for many of the reasons that have been discussed in previous episodes. If you work at a big company with some monolithic characteristics, this episode might give you some good arguments to bring to your manager about why and how to move to microservices.

Ep 303Serverless Code with Ryan Scott Brown
The unit of computation has evolved from on premise servers to virtual machines in the cloud to containers running in those virtual machines. Serverless computation is another stage in the evolution of computational unit management. With a serverless architecture, a function call to the cloud spins up a transient container, calls the function on that container, and then spins down the container. Ryan Scott Brown joins the show today to discuss the benefits and consequences of serverless computing. With containers and VMs, we still have to worry that the resources we are spinning up in the cloud will run without being utilized. Serverless computing gives us more control over these compute resources, so that we don’t have unused servers that we are paying for.

Ep 302Algorithm Marketplace with Diego Oppenheimer of Algorithmia
Algorithmia is marketplace for algorithms. A software engineer who writes an algorithm for image processing or spam detection or TF-IDF can turn that algorithm into a RESTful API to be consumed by other developers. Different algorithms can be composed together to build even higher level applications. Diego Oppenheimer is the CEO of Algorithmia, and he joins the show today to explain how Algorithmia works. The company has developed its own container orchestration and management service, and operates similarly to the serverless computing paradigms that we have discussed on recent episodes of Software Engineering Daily. Diego also talks about the marketplace dynamics of building a platform for developers to sell algorithms to each other.

Ep 301Internet of Things with Azure’s Steve Busby
The Internet of Things is becoming a reality. Factories are being outfitted with sensors, temperature monitors, and other data gathering devices. In agriculture, farms are becoming more efficient thanks to soil monitoring devices and automated pesticide regulation. In our homes, refrigerators, alarm clocks, and mirrors are becoming “smart”. Steve Busby joins the show today to talk about the big picture: how the Internet of things works, from data ingestion to processing to feedback. Steve works at Microsoft in the Azure IoT division, and he discusses the problems which companies are having and the solutions that are available today–and where we are going in the future.

Ep 300Secret Management and Vault with Hashicorp’s Seth Vargo
Every software application has secrets. User passwords and database credentials must be managed carefully, because poor access controls can lead to disaster scenarios. Vault is a tool for secret management, developed at Hashicorp, a company that builds software tools for application delivery and infrastructure management. Seth Vargo is a software engineer and open source advocate at Hashicorp, and in today’s episode he discusses the advantages of having a single tool to manage all of your secrets. If you aren’t a security expert, don’t worry, we discuss some of the basics of security. And if you are a security expert, you will appreciate the comparisons we discuss between Hashicorp and other tools that have been used for secret management.

Ep 299Google’s Site Reliability Engineering with Todd Underwood
Google’s site reliability engineers are responsible for maintaining the highly available services that power the Google software that we all use on a regular basis. O’Reilly recently published the book “Site Reliability Engineering: How Google Runs Production Systems”, and the book provides a comprehensive window into how the site reliability engineering role works. Todd Underwood is a director of site reliability engineering. On today’s episode, Todd explains how the role of a SRE relates to devops. We discuss the relationship between the engineers who are developing Google services, and the SREs who are maintaining it. Google’s internal data center operating system “Borg” is also discussed.

Ep 298Female Pursuit of Computer Science with Jennifer Wang
Google researcher Jennifer Wang co-wrote a paper called “Gender Differences in Factors Influencing Pursuit of Computer Science and Related Fields”. The paper focuses on a survey of 1700 high school and college students, and takes a statistical approach to understanding why women are not pursuing computer science. In our conversation, Jennifer talks about the two influences that lead to fewer women in computer science: encouragement and exposure. The problem of encouragement: women often do not receive encouragement to go into computer science. The problem of exposure: women are often unaware that computer science even exists. On this episode, we explore the roots of these problems, and other results of her demographic study of young students.

Ep 297JavaScript Concurrency with Kyle Simpson
JavaScript programming usually is done through the use of frameworks, such as ReactJS, AngularJS, and EmberJS. These frameworks abstract away some of the messy details of JavaScript, and simplify web development so that engineers can build products at a faster pace. When we build software using JavaScript frameworks, we are missing out on some of the richness of the JavaScript language itself. Kyle Simpson is the author of “You Don’t Know JS”, a series of books that suggests that JavaScript developers should start from the ground up, not from the top down. By learning the basics of JavaScript, a software engineer can learn the timeless fundamentals that will not disappear with the creation of next week’s hottest framework. After exploring the idea of frameworks versus raw JavaScript, Kyle and I discuss asynchronous JavaScript, from concurrency to the observer pattern. Sponsors
Music

Ep 295Serverless Framework with Austen Collins
Virtual machines were the unit of cloud computation for many years. Amazon Web Services pioneered the democratized model of allowing anyone to deploy a service to the cloud, running on a virtual machine on Amazon’s servers. After virtual machines, containers have become the unit of scale in the cloud. We break up our virtualized servers into even smaller units of computation called containers. Today, the unit of compute is getting reduced even more, with the introduction of serverless architecture. Serverless architectures started getting talked about after Amazon Web Services released a service called AWS Lambda, which allows users to have pieces of code run in response to events. Programmers write a function and hand it off to Amazon, and Amazon will run that function call, and only charge the programmer when the function is actually called. This is in contrast to the cost model of containers or virtual machines, which users pay for even while they are running idle. Today’s guest Austen Collins believes that serverless computing is the model of the future, and he created a company called Serverless around this idea. His company Serverless provides an application framework for building applications exclusively on Amazon Web Services Lambda.

Ep 294Management and Hiring with Jon Emerson
Engineering managers start out as engineers. Eventually, there is a fork in their career road where an engineer can choose to move up into management or continue on as an engineer in a more senior role. Changing to management involves an increase in responsibilities, a different set of goals to focus on. Jon Emerson was working at Google as an engineer when a project he was working on started to get more attention. He moved into management, and spent several years at Google as a manager. Today, he works at Hired as a director of engineering, leading four different teams. Full disclosure: Hired is a sponsor of Software Engineering Daily, but this episode is not about Hired itself. We do start out with a discussion of the hiring process, because of Jon’s domain expertise around hiring, but most of our conversation focuses on the role of a manager, and the role of a director. Most of the episodes of Software Engineering Daily focus on the day-to-day life of an engineer, so it was interesting to get a perspective from someone higher up the management chain, who has more visibility over entire software projects.

Ep 293Phone Spam with Truecaller CTO Umut Alp
The war against spam has been going on for decades. Email spam blockers and ad blockers help protect us from unwanted messages in our communication and browsing experience. These spam prevention tools are powered by machine learning, which catches most of the emails and ads that we don’t want to see. TrueCaller is a company that is bringing this quality of spam detection to our phone call systems. Umut Alp is the CTO of TrueCaller, and he joins the show today to break down the engineering problems of preventing telephone call spam. Users of TrueCaller install it on their phone, and the software allows users to report when they have received a spam call. Using this reporting mechanism, and other learning algorithms, TrueCaller is able to learn what types of calls it should block from being accepted by your phone. Today on Software Engineering Daily, we discuss cell phone spam prevention.

Ep 292Tech Girls Movement with Jenine Beekhuyzen
The software industry has a severe lack of women. There are numerous root causes of this diversity problem. Families do not encourage women to enter math and science. The media portrays most programmers as white males. Our industry often picks up on the signals of the broader society and perpetuates them. Reversing this trend of low female involvement in computer science could have tremendous positive impact, and Jenine Beekhuyzen joins the show today to discuss how she she is taking action to improve the pipeline of young women entering computer science.

Ep 291Google’s Polymer Project with Rob Dodson
Smart phone apps have better performance than web apps. When we have an application that we use on a regular basis, we download that application to a smart phone rather than using the browser based version on our mobile browser. Google’s Polymer Project wants to improve the gap between native app performance and mobile web app performance. The key problem with mobile web is that we are sending huge JavaScript bundles to mobile devices, which inhibits performance. The Polymer Project is working to build more functionality into our mobile browsers so that it is easier to load these heavy web applications. Rob Dodson is a developer advocate with Google. Today’s episode explores the past, present, and future of web application development, from jQuery to React to progressive web apps. The Polymer project also represents a push towards better support for people in developing countries, where internet connections are less reliable. Sponsors

Ep 290Software Editorialism with Practical Dev’s Ben Halpern
Most programmers spend lots of their time reading content about software. Since our field changes so rapidly, engineers consume news and editorials voraciously, trying to keep up with the impossibly fast pace. The Practical Dev is a collection of blog posts, editorials, and interviews that was created to help with that end. Ben Halpern is the creator of Practical Dev, and he joins the show to discuss software editorialism. The goal of the Practical Dev to help developers grow and learn, and Ben is working towards that goal by providing a platform for engineers to write long-form content.

Ep 289Scaling PostgreSQL with Citus Data’s Ozgun Erdogan
Ten years ago, databases were much simpler. Most companies would only have one or two types of databases in production. Today, the age of one-size-fits-all is over. Companies have multiple databases to deal with different types of use cases, and databases have become distributed to multiple nodes in order to be scalable. Ozgun Erdogan of Citus Data joins the show to give us a modern look at databases. Ozgun suggests that PostgreSQL alone can perform most of the work that we are trying to get from our variety of databases. We discuss how Citus Data scales Postgres, and Ozgun contrasts an all-Postgres architecture with other types of databases such as the “NewSQL” class of databases, like MemSQL and VoltDB.

Ep 288Kubernetes and OpenShift with Clayton Coleman
Kubernetes is the container management platform that came out of Google’s experiences managing data centers. Kubernetes abstracts away many of the frustrations of distributed systems management. OpenShift is a platform built on top of Kubernetes to provide an additional layer of usability. Clayton Coleman is the lead engineer of OpenShift, and in our conversation today, we start with the basics of Kubernetes, then talk about OpenShift–Clayton explains why we need another abstraction on top of Kubernetes. Near the end of the conversation, we discuss the current state of cloud products–which can be confusing. Mesosphere, Swarm, Kubernetes, OpenStack, ECS–why do we need all of these different products? Clayton gives his perspective, and explains why it is not going to get any less confusing any time soon.

Ep 287Boot Camps, Mesosphere, and Open-Source with Kenny Tran
Coding boot camps are a subject of controversy. Critics of boot camps defend the conventional university system, and argue that boot camp graduates do not have enough experience to write quality software. But the reality is that some boot camp graduates have found success from this new educational path. After graduating high school, Kenny Tran attended one coding boot camp, then spent some time living at home absorbed in his personal projects. Eventually, he went to second coding boot camp–his form of graduate school. During his second boot camp, Kenny worked on PurifyCSS, a module that can reduce the size of front-end projects by 60%. Today, Kenny works at the groundbreaking company Mesosphere–completing a career arc that proves coding boot camps are enough of an education to rival traditional university learning.

Ep 286Infrastructure as Code with SaltStack’s David Boucha
Infrastructure-as-code is a trend that has been popularized over the past decade, as cloud computing and distributed systems have become a part of every technology company. Tools like Salt, Puppet, Chef, and Ansible allow us to manage servers and processes from the command line. David Boucha works at SaltStack, the company that makes Salt. Salt is a platform that provides configuration management and remote execution. In our conversation today, we discuss how the distributed systems architecture of Salt works, and how it can be used in practice. We also discuss infrastructure-as-code in the context of modern infrastructure tools like containers and Kubernetes.

Ep 285Solar Investment and Architectural Strategy at Wunder Capital
Solar energy is a growing market. Improvements in hardware have led to some people predicting that solar energy will be powering the world within the next few decades. Undoubtedly, a large percentage of our current energy infrastructure will be replaced by solar in the near future. Replacing our old, inefficient power grid requires massive investment. On Software Engineering Daily, the resources we usually talk about are memory, bandwidth, storage, and other computer science concepts. David Reiss joins us today to talk about energy, which is a resource so fundamental that we usually don’t even consider it. David is the CTO of Wunder Capital, a fintech company that facilitates investments in solar power. Much of this conversation centers on the economics of the solar energy market, but we also discuss software topics–machine learning and why Wunder Capital’s software architecture is a monolith by design today. Sponsors

Ep 284Minecraft Programming with Gabriel Simmer
Minecraft is a sandbox video game in which players build constructions out of 3-D cubes in a procedurally generated world. Minecraft is the best-selling PC game of all time. But Minecraft is not just a game. It is a platform for creativity, used by players within the game as well as programmers outside of it. Gabriel Simmer is a 16-year-old programmer who build NodeMC, a tool that wraps around the Minecraft server process. NodeMC can be used to build dashboards for Minecraft, and spin up additional Minecraft servers. In our conversation, Gabriel explains why people are so excited about Minecraft, how people are hacking Minecraft, and what the future of Minecraft is now that Microsoft has acquired it.

Ep 283Rust with Steve Klabnik
Rust is a systems programming language being developed at Mozilla. Rust has features of a high-level functional language like Scala and a low-level, performance-driven language like C++. Steve Klabnik is a developer program member with Mozilla. In this episode, he discusses how Rust looks at memory management, type safety, mutability, and concurrency. We also dive into a discussion of the low level virtual machine, also known as the LLVM, which the language Swift is built on.

Ep 282Kafka, Storm, and Cassandra: Keen IO’s Analytics Architecture with Dan Kador
The process of building a software project requires us to make so many architectural decisions. Which programming languages should be used? Which cloud service provider? Which database? A newer type of building block is the analytics platform. Companies need to track events, aggregate metrics, and change the user’s experience based on aggregated data. Dan Kador is a co-founder of Keen IO, and he joins us to discuss the rise of analytics platforms. Keen’s architecture is based on Storm, Cassandra, and Kafka, and Keen uses these different building blocks to create a scalable, reliable analytics backend. On today’s episode, we explore the usage of analytics, the architecture of Keen’s backend system, and the business model of an analytics as a service company.

Ep 281Erlang Systems Design with Francesco Cesarini
Erlang is a programming language with primitives that help software engineers build distributed systems. When a process is malfunctioning in Erlang, the philosophy of the language is to let the process crash–and in a distributed system where unexpected faults happen on a regular basis, this philosophy of “let it crash” simplifies how we reason about an Erlang system. Other distributed systems advantages of Erlang include the garbage collection strategy. Each process in Erlang has its own garbage collector which means makes it easier to construct systems without a stop-the-world garbage collection. Francesco Cesarini is the founder of Erlang Solutions, and he joins the show today to discuss the book that he wrote with Steve Vinoski, called Designing for Scalability with Erlang/OTP.

Ep 280Google’s Microservices: Kubernetes and gRPC with Sandeep Dinesh
Google has built a microservices architecture on top of the internal container management system called Borg. These services communicate over an internal protocol known as Stubby. Borg and Stubby are tightly coupled to Google’s infrastructure–it would not make sense for Google to open source them–but Google has worked with the open source community to develop open source projects with the core functionality of Borg and Stubby. These projects are known as Kubernetes and gRPC. Sandeep Dinesh works on the Google Cloud Platform team as a developer advocate. Our conversation explores how a client application request from an app like Gmail would communicate with Google’s servers, where the request is handled by a network of microservices. We also talk about where Google Cloud Platform is evolving, and how it offers a competitive, differentiated model from Amazon Web Services.

Ep 279Netflix’s Data Pipeline with Steven Wu
At Netflix, 500 billion events and 1.3 petabytes of data are ingested by the system per day. This includes video viewing activities, error logs, and performance events. On today’s episode, we dive deep into the data pipeline of Netflix, and how it evolved from their 1.0 version to the modern 2.0 version. Before listening to this episode, check out the blog post that inspired it.

Ep 278Dropbox’s Magic Pocket with James Cowling
Dropbox has been storing files on Amazon Web Services for 8 years, and Dropbox’s core business is storing files. For the past three years, Dropbox has been working on a project to migrate its file storage from Amazon Web Services to its own custom built infrastructure. Magic Pocket is the name of Dropbox’s new infrastructure layer, and it gives Dropbox more control and improved economics. James Cowling leads the storage team at Dropbox. On today’s episode, James takes us into the architecture of Dropbox and explains how the team moved all of the user file storage from Amazon S3 to Dropbox’s Magic Pocket infrastructure. Dropbox’s architecture is built with a focus on simplicity–and there are numerous challenges to maintaining that simplicity in the face of an extremely complex problem like this.

Ep 277Decentralization: Ethereum, Bitcoin, and IPFS with Karl Floersh
Almost a year ago, Software Engineering Daily aired a week of shows about decentralized technologies like Bitcoin, Ethereum, and IPFS. Bitcoin has established itself as a stable network, but it can only be used for financial transactions. Ethereum is a global computer built on a blockchain, but it does not have the adoption of Bitcoin. IPFS is a distributed data store with an incentivization layer where you can pay strangers to store your content. Today’s guest is Karl Floersh, an engineer working on decentralized technology. Karl has written several posts about how to build decentralized applications, as well as how the future might look once these decentralized technologies have gotten traction.

Ep 276Distributed Systems Tradeoffs with Camille Fournier
Distributed systems products are often marketed with terms like “real-time data” and “hassle-free scaling”, but what do those terms actually mean? Is data in a distributed system ever reliably “real time”? Do we ever have strong enough plans about our scalability strategy to say that scaling will be “hassle free”? Camille Fournier joins us today to discuss distributed systems in practice. Like everything in else in computer science, distributed systems are all about tradeoffs–and picking the right sets of tradeoffs in our distributed system will affect the entire organization that is building that system. We also discuss the Cloud Native Computing Foundation, which is similar to the Apache Foundation, but specifically for cloud technologies. The CNCF is likely to have strong impact on the way we build software for a long time to come.

Ep 275Crate.io and Distributed SQL with Jodok Batlogg
Distributed databases are difficult to operate, and Crate.io wants to change that. Crate is a fast, scalable, easy-to-use SQL database that is built to run in containerized environments. An average software company runs several databases–MySQL for relational store, MongoDB for a document database, HDFS for blob storage and data warehouse, elastic search for search. On today’s show, Jodok Batlogg from Crate discuss the distributed architecture of Crate, and breaks down the use cases, from Microservices to data warehousing.

Ep 274Azure Stream Analytics with Santosh Balasubramanian
Microsoft has built a suite of technologies on top of its Azure infrastructure as a service. Today, we discuss Azure Stream Analytics, a real-time event processing engine developed at Microsoft. Azure Streaming allows for constant querying of incoming data streams, and my guest Santosh Balasubramanian discusses Azure and the movement from batch processing to stream processing.

Ep 273Spark and Cassandra with Tim Berglund
Apache Spark is a framework for fast, distributed, in-memory analysis. Apache Cassandra is a distributed database management system that provides high availability and fast throughput. Today, we are collecting fast, big data streams from user behavior, smart phones and sensors, and the disk checkpointing of and query language of Hadoop MapReduce is no longer adequate. Tim Berglund from Datastax came on Software Engineering Daily to explain how Apache Cassandra in a popular episode a few weeks ago. On this episode, Tim returns to discuss how Spark and Cassandra can be used together to provide a stack with the analytics and storage we need for today’s distributed computing environment.