June 7th, 2013

Chocolate Cake

RSS icon RSS Category: Uncategorized
Fallback Featured Image

Chocolate Cake (Wednesday, June 5, 2013) 

You know how sometimes you have one bite of really good chocolate cake, or a really amazing peach and totally assume that you could eat another 30lbs of whatever without regard for good manners or physical limitations?  Yeah. Decreasing marginal returns dictate that it almost always turns out that the last bite isn’t as good as the first one – having a little and having a lot are different.

Similarly, ingesting 1000 bytes of data and 1 byte are pretty different, and when you’re used to little bytes and start fooling around with the big ones the differences might not be immediately obvious or intuitive (maybe they are, and if that’s the case – awesome! Now go eat your cake).

When I’m trying to make sense of a problem I like to start with a small example and work through it to get a feel for the mechanics. With Big Data this gets a little weird, since we’re almost always mining, so we don’t always know well what to look for or expect, and because we need some intuition for how to get from the small to the big.  To help that, I am trying to build some intuitive explanations.  You can look at them topically under the posts beginning with header “Big vs. Little…”

Sometimes it is the case that using H2O to look at small data sets really makes no sense for whatever reason. In those cases we’ll talk about why, and I’ll use R for comparison. I’ll also provide you with relevant output for each (so that you can see how to get from one to the other). If you’re not familiar with R go here.  Additionally, it’s worth mentioning that I’m tackling one set of assumptions at a time, so in general I’ll work as though we are going through some ad-hoc analysis instead of post-hoc analysis.  There are some super cool differences between mucking vs. mining, but I want to talk about those separately.

Leave a Reply

AI-Driven Predictive Maintenance with H2O Hybrid Cloud

According to a study conducted by Wall Street Journal, unplanned downtime costs industrial manufacturers an

August 2, 2021 - by Parul Pandey
What are we buying today?

Note: this is a guest blog post by Shrinidhi Narasimhan. It’s 2021 and recommendation engines are

July 5, 2021 - by Rohan Rao
The Emergence of Automated Machine Learning in Industry

This post was originally published by K-Tech, Centre of Excellence for Data Science and AI,

June 30, 2021 - by Parul Pandey
What does it take to win a Kaggle competition? Let’s hear it from the winner himself.

In this series of interviews, I present the stories of established Data Scientists and Kaggle

June 14, 2021 - by Parul Pandey
Snowflake on H2O.ai
H2O Integrates with Snowflake Snowpark/Java UDFs: How to better leverage the Snowflake Data Marketplace and deploy In-Database

One of the goals of machine learning is to find unknown predictive features, even hidden

June 9, 2021 - by Eric Gudgion
Getting the best out of H2O.ai’s academic program

“H2O.ai provides impressively scalable implementations of many of the important machine learning tools in a

May 19, 2021 - by Ana Visneski and Jo-Fai Chow

Start your 14-day free trial today