The DataOps Blog
Where Change Is Welcome
The Next Chapter for StreamSets
Arvind Prabhakar and I co-founded StreamSets in 2014 with an audacious vision: data should be the lifeblood of the enterprise.…
Continuous Ingest to Elasticsearch (video)
You can send log files to Elasticsearch using StreamSets Data Collector. Watch this short step-by-step to see how.
Continuous Ingest in the Face of Data Drift – Part 2 (from the Cloudera Vision Blog)
In my previous post I discussed the causes and impacts of data drift, a natural consequence of Big Data which creates serious data quality and data pipeline operational issues. Now I will describe the features of StreamSets Data Collector, how they address ingesting data in a “drifty” environment and describe some common use cases. StreamSets was founded to deliver a…
Continuous Ingest in the Face of Data Drift (from the Cloudera Vision Blog)
Big data has come a long way, with adoption accelerating as CIOs recognize the business value of extracting insights from the troves of data collected by their companies and business partners. But, as is often the case with innovations, mainstream adoption of big data has exposed a new challenge: how to ingest data continuously from any source and with high…
Announcing StreamSets Data Collector 1.2.0.0
We are very excited to announce the next version of the StreamSets Data Collector. This version has seen over 250 JIRAs with a host of new features, performance enhancements and bug fixes. Without further ado, here's what's new in 1.2.0.0: Versioning Scheme To keep up with an incredible volume of feature requests from customers and the community, we are introducing…
StreamSets Monitoring with Grafana, InfluxDB, and jmxtrans
The ability to monitor your critical infrastructure is a must, and we designed the StreamSets Data Collector (SDC) with this in mind: metrics are exposed through both the REST API and JMX. While there are many approaches to monitoring these metrics, let’s walk through a specific end-to-end example using jmxtrans to collect metrics, InfluxDB to store them, and Grafana to…