BT

Apache Spark 2.0 Technical Preview

by Alex Giamas on  May 31, 2016

Two years after the first release of Apache Spark, Databricks announced the technical preview of Apache Spark 2.0 , based on upstream branch 2.0.0-preview. The preview is not ready for production, neither in terms of stability nor API, but is a release intended to gather feedback from the community ahead of the general availability of the release.

Amazon Releases Kinesis Service Update

by Kent Weare on  May 23, 2016

Amazon has recently announced an update to their Amazon Kinesis Service. In this update, three new features have been added to Amazon Kinesis Streams and Amazon Kinesis Firehose including support for Elasticsearch Service Integration, Shard-Level Metrics and Time-Based Iterators.

Precision Medicine Modeling Demonstration with Spark on EMR, ADAM, and the 1000 Genomes Project

by Dylan Raithel on  May 19, 2016

AWS engineers Christopher Crosbie and Ujjwal Ratan detail using Spark on EMR for precision medicine data analysis on the ADAM platform with data from the 1000 genomes project.

The Broad Institute Migrates Genome Sequencing Pipeline to Google Cloud Platform

by Dylan Raithel on  May 13, 2016

Genomic data sequencing and subsequent analysis faces large data volume challenges that several organizations are solving with cloud services. The Broad Institute detailed their experience with petabyte scale sequencing pipelines last month through the Google Research Blog and is detailed here by InfoQ.

Deep Mind Discloses Details to InfoQ about NHS Partnership amid Reports of Vast Patient Data Access

by Dylan Raithel on  May 05, 2016 1

After months of awaiting details about the NHS and Google DeepMind partnership InfoQ gains insights into recent claims of widespread patient data access.

Elephant in the Cloud - Hadoop as a Service

by Srini Penchikala on  May 02, 2016 2

Hadoop and other big data technologies revolutionized the way organizations run data analytics but the organizations are still facing challenges with operating costs of using these technologies for on-premise data processing. Ashish Thusoo recently spoke at Enterprise Data World Conference about Hadoop as a service offering that helps organizations bridge the gaps with these capabilities.

AirFlow Joins Apache Incubator

by Alex Giamas on  Apr 29, 2016

AirFlow recently joined the Apache Incubator program. AirFlow is a workflow and scheduling system designed to manage data pipelines. Developed by AirBnb for their internal usage, it was open sourced last September, as previously reported by InfoQ.

Operational Data Stream and Batch Processing at Netflix with Mantis

by Dylan Raithel on  Apr 27, 2016

Operational Data Stream and Batch Processing at Netflix with Mantis

Neo4j 3.0 Released with Binary Communication Protocol and Standardised Drivers

by Alex Blewitt on  Apr 26, 2016

Today at GraphConnect Europe 2016, Neo Technology announced the release of Neo4j 3.0, which includes a new binary protocol for transmitting data between server and client, and a new set of standardised drivers for interacting with the database, along with stored procedure support and higher performance and capacity. InfoQ spoke to Neo Technology to find out more.

Google Cloud Machine Learning and Tensor Flow Alpha Release

by Dylan Raithel on  Apr 18, 2016

Late last month Google released an alpha version of their TensorFlow (TF) integrated cloud machine learning service as a response to a growing need to make their Tensor Flow library to run at scale on the Google Cloud Platform (GCP). Google describes several new feature sets around making TF usage scale by integrating several pieces of the GCP like Dataproc, a managed Hadoop and Spark service.

Apache Flink 1.0.0 is Released

by Rags Srinivas on  Mar 24, 2016

InfoQ's Rags Srinivas caught up with Stephan Ewen, a project committer for Apache Flink about the 1.0.0 Release and the roadmap

Databricks Integrates Spark and TensorFlow for Deep Learning

by Dylan Raithel on  Mar 12, 2016

Since announcements late last year about Google open-sourcing TensorFlow, the company’s open-source library for machine learning, and previous coverage at InfoQ, the data-science community has had an opportunity to try out TensorFlow for their own projects.

Funnel Analysis at Twitter for Improving User Engagement

by Srini Penchikala on  Feb 25, 2016

Funnel analysis is used to analyze a sequence of events to help with user engagement on a website or a mobile application. Data Science team at Twitter uses this concept to learn how users interact with user interfaces during sign up or tweeting for improving user engagement with Twitter.

AlphaGo: Google and DeepMind Publish Seminal AI Work

by Dylan Raithel on  Feb 23, 2016

A game simulation at Google's Deep Mind defeated expert humans at Go last month in a breakthrough for AI. Go is considered one of the great unsolved problems in AI.

Benchmarking Netflix Dynomite with Redis on AWS

by Alex Giamas on  Feb 03, 2016

Last year, Netflix Cloud Database Engineering (CDE) team introduced Dynomite. Dynomite is a proxy layer, aiming to turn any non-distributed database into a sharded, multi-region replication aware distributed database system. Now Netflix released a benchmark using Dynomite with Redis in AWS infrastructure.

General Feedback
Bugs
Advertising
Editorial
Marketing
InfoQ.com and all content copyright © 2006-2016 C4Media Inc. InfoQ.com hosted at Contegix, the best ISP we've ever worked with.
Privacy policy
BT