Skip to content
TrackPodcasts
technologyJun 2, 202252:30pending

Apache Spark (Pt. 2): MLlib - ML 074

About this episode

MLlib is Apache Spark's scalable machine learning library. Today, Ben and Michael discuss the ease of use, performance, algorithms, and utilities included in this library and how to execute the best ML workflow with MLlib.

In this episode...

  1. Why stick with Spark libraries vs. a single node operation?
  2. What algorithms are not in Spark Lib?
  3. What is the min. package set to use for supervised learning?
  4. Modeling and validation
  5. Down-sampling your data
  6. MLlib vs. scikit-learn
  7. Resources

Sponsors


Links



Advertising Inquiries: https://redcircle.com/brands

Privacy & Opt-Out: https://redcircle.com/privacy

Become a supporter of this podcast: https://www.spreaker.com/podcast/adventures-in-machine-learning--6102041/support.

Get every episode summarized

Each time Adventures in Machine Learning publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

Apache Spark (Pt. 2): MLlib - ML 074

Adventures in Machine Learning

0:00
52:30

More episodes

More from Adventures in Machine Learning

View all episodes →