InfoQ Homepage Big Data Content on InfoQ

News

RSS Feed

Newer Older

Apache Spark 1.3 Released, Data Frames, Spark SQL, and MLlib Improvements

Apache Spark has released version 1.3 of their project. The main improvements are the addition of the DataFrames API, better maturity of the Spark SQL, as well as a number of new methods added to the machine learning library MLlib, and better integration of Spark Streaming with Apache Kafka.

Mikio Braun
on Mar 16, 2015
MongoDB 3.0 - WiredTiger Storage Engine and Updated MMS

Some time ago, when MongoDB 2.6 was released Kelly Stirman, Director of Products at MongoDB answered our questions regarding the latest release. Now with MongoDB 3.0 announced for March and MongoDB 3.0 RC-8 already available, it’s time to see in more detail what WiredTiger storage engine, new and improved MMS and storage compression can bring to NoSQL users.

Alex Giamas
on Feb 20, 2015
Pivotal Open Sources Their Big Data Suite

Pivotal has decided to open source core components of their Big Data Suite and has announced the Open Data Platform, an initiative promoting open source and standardization for Big Data.

Abel Avram
on Feb 19, 2015
Project Pachyderm Aims to Build a "Modern" Hadoop on Docker

Project Pachyderm Aims to Build "Modern" Hadoop using Docker and CoreOS.

Matt Kapilevich
on Feb 17, 2015
Apache Hive 1.0 Released, HiveServer2 Becomes Main Engine, Stable API Defined

Apache Hive has released version 1.0 of their project on February 6th, 2015. Originally planned as version 0.14.1, the community voted to change the version numbering to 1.0.0 to reflect the amount of maturity the project has reached.

Mikio Braun
on Feb 11, 2015
EMRFS Brings Consistency to Amazon S3

Amazon recently announced EMRFS, an implementation of HDFS that allows EMR clusters to use S3 with a stronger consistency model. When enabled, this new feature keeps track of operations performed on S3 and provides list consistency, delete consistency and read-after-write-consistency, for any cluster created with Amazon Machine Image (AMI) version 3.2.1 or greater.

Jérôme Serrano
on Jan 27, 2015
Apache Flink 0.8.0 Released, Roadmap for 2015 Published

Apache Flink has released the version 0.8.0 of their project. Besides the usual performance, compatibility, and stability improvements, it has also added a streaming Scala API, where streaming capabilities had so far been missing. Apache Flink has also been promoted to the top-level of the Apache projects recently after joining the incubator roughly nine months ago.

Mikio Braun
on Jan 22, 2015
Google on the Technical Debt of Machine Learning

A number of Google researchers and engineers presented their view on the technical debt of using machine learning at a NIPS workshop. They identified different aspects of technical debt and came to the conclusion that without proper care, using machine learning or complex data analysis in your company can induce new kinds of technical debt different from classical software engineering.

Mikio Braun
on Jan 15, 2015
Apache Spark 1.2.0 Supports Netty-based Implementation, High Availability and Machine Learning APIs

Apache Spark 1.2.0 was released with Netty-based implementation, High Availability and Machine Learning APIs. It represents the work of 172 contributors from over 60 institutions and comprises more than 1000 patches. InfoQ talks with Patrick Wendell, a Spark committer and PMC member.

Rags Srinivas
on Jan 07, 2015
Splunk Enterprise 6.2 Supports Instant Pivot and Enhanced Event Pattern Detection

The latest version of big data analytics tools Splunk Enterprise and Hunk support instant pivot, enhanced event pattern detection, and prebuilt dashboard panels. Splunk Inc., provider of the software platform for operational intelligence, recently announced the general availability (GA) of version 6.2 of Splunk Enterprise and Hunk: Splunk Analytics for Hadoop and NoSQL Data Stores.

Srini Penchikala
on Dec 21, 2014
New and Interesting on ThoughtWorks Radar Jan 2015

ThoughtWorks has published a digital preview of the January 2015 radar, providing opinion on techniques, tools, platforms and languages and taking a snapshot of the current trends in software technology.

Abel Avram
on Dec 19, 2014
Splice Machine Version 1.0 Supports Integration with Hadoop and Analytic Window Functions

Splice Machine version 1.0 supports analytic window functions and integration with Hadoop ecosystem. Splice Machine team recently released their Hadoop based RDBMS data management solution that can be used for transactional workloads on Hadoop.

Srini Penchikala
on Dec 18, 2014
Google Open Sources Cloud Dataflow Java SDK

Google announced earlier this year their Cloud Dataflow, a service and SDK for processing large amounts of data in batches or real time. Now they have open sourced the Dataflow Java SDK, enabling developers to see how it works and possibly use the SDK for services running on-premises or in other clouds.

Abel Avram
on Dec 18, 2014
LinkedIn Open Sources Cubert With an Eye To Big Data Analytics

LinkedIn recently open sourced Cubert, its High Performance Computation Engine for Complex Big Data Analytics. Cubert is a framework written for analysts and data scientists in mind.Developed completely in Java and expressed as a scripting language, Cubert is designed for complex joins and aggregations that frequently arise in the reporting world.

Alex Giamas
on Dec 17, 2014
Agile View of Big Data

An agile view of Big Data, wherein data is viewed as a real time stream, offers a new look at how data is managed. Using an agile data infrastructure, organizations can conquer Big Data challenges with a level of ease, flexibility and performance. White paper by codeFutures describes the Agile view of Big Data.

Savita Pahuja
on Dec 16, 2014

Newer News

Older News

InfoQ Software Architects' Newsletter

News