June 26th, 2014

Learn to manage, munge, and model big data with H2O on the Hortonworks Sandbox

RSS icon RSS Category: Uncategorized

Working with big data might seem like a daunting task if like me, you’ve spent the majority of your college years doing pencil and paper proofs. Big data for me was anything that took longer than 30 minutes to ingest into single threaded R.
For mathematicians and statisticians looking to understand widely used data platforms like Hadoop for data storage and data management, Hortonworks Sandbox is an awesome all-in-one self-teaching tool. Getting a standalone Hadoop environment on your personal computer is as easy as launching a VM.

To actually start doing predictive analytics, launch H2O on the server either as a simple JVM or a mapper task that’ll utilize all the nodes in the cluster. When it comes time to actually move from a test and research setting to a production one, the same installation and launch holds for however many nodes you add to the cluster.
H2O and Hortonworks Sandbox will turn the uninitiated into data scientists with a gentle sloping learning curve. Both H2O and Hortonworks are open source big data powerhouses that you can learn from and perhaps eventually contribute to.
It’s free to try so get started now with the following tutorial : Predictive Analytics on H2O and Hortonworks Data Platform

For more information about how H2O operates on Hadoop check out : H2O on Hadoop

Leave a Reply

AI-Driven Predictive Maintenance with H2O Hybrid Cloud

According to a study conducted by Wall Street Journal, unplanned downtime costs industrial manufacturers an

August 2, 2021 - by Parul Pandey
What are we buying today?

Note: this is a guest blog post by Shrinidhi Narasimhan. It’s 2021 and recommendation engines are

July 5, 2021 - by Rohan Rao
The Emergence of Automated Machine Learning in Industry

This post was originally published by K-Tech, Centre of Excellence for Data Science and AI,

June 30, 2021 - by Parul Pandey
What does it take to win a Kaggle competition? Let’s hear it from the winner himself.

In this series of interviews, I present the stories of established Data Scientists and Kaggle

June 14, 2021 - by Parul Pandey
Snowflake on H2O.ai
H2O Integrates with Snowflake Snowpark/Java UDFs: How to better leverage the Snowflake Data Marketplace and deploy In-Database

One of the goals of machine learning is to find unknown predictive features, even hidden

June 9, 2021 - by Eric Gudgion
Getting the best out of H2O.ai’s academic program

“H2O.ai provides impressively scalable implementations of many of the important machine learning tools in a

May 19, 2021 - by Ana Visneski and Jo-Fai Chow

Start your 14-day free trial today