Brooklyn Park , MN
|Remote
|Contract
Brooklyn Park, MN
|Remote
|Contract
Join a high-impact engineering team supporting a large-scale Hadoop platform for enterprise analytics and data processing. This role focuses on keeping a complex data environment reliable, performant, and observable while partnering with cross-functional teams to improve stability, automation, and operational excellence.
Responsibilities
- Design, maintain, and optimize a scalable Hadoop platform that supports critical data processing and analytics workloads.
- Troubleshoot and resolve performance issues across core Hadoop ecosystem tools and services, including HDFS, YARN, MapReduce, Hive, Spark, and HBase.
- Build and standardize monitoring, health checks, and alerting to improve platform visibility and reduce downtime.
- Use Linux expertise, including Ubuntu, to investigate system bottlenecks, tune kernel settings, and improve resource usage.
- Partner with data engineering, infrastructure, and platform teams to improve job efficiency and cluster stability.
- Establish operational best practices for capacity planning, performance tuning, and version upgrades.
- Automate repetitive tasks with scripting tools such as Python, Bash, or Scala to improve workflow efficiency.
- Support data security, governance, and compliance practices across the Hadoop environment.
- Use automation frameworks such as Chef and/or Ansible to manage large-scale cluster operations.
- Lead root cause analysis for recurring incidents and help implement preventive solutions.
- Document runbooks, standards, and configuration guidance to support consistent platform management and knowledge sharing.
Skills
- Hands-on experience with Apache Hadoop and core ecosystem components.
- Strong working knowledge of HDFS, Hive, Spark, and Java.
- Solid Linux administration skills, with experience troubleshooting and tuning Ubuntu environments.
- Experience supporting large-scale data platforms and diagnosing cluster-level issues.
- Ability to implement monitoring, observability, and automated alerting solutions.
- Experience with configuration management tools such as Chef and/or Ansible.
- Ability to automate operational work using scripting languages.
- Strong analytical and problem-solving skills with a proactive approach to incident prevention.
- Clear communication skills and the ability to document technical processes effectively.
Preferred Skills
- Experience with Grafana for dashboards and platform visibility.
- Working knowledge of Python.
- Exposure to Trino.
- Background supporting enterprise data platform operations in a collaborative environment.
- Familiarity with best practices for performance tuning, operational resilience, and platform governance.
Horizontal is committed to fostering an inclusive and welcoming environment where different perspectives are valued and every team member can contribute meaningfully. We encourage candidates from all backgrounds to apply and support diversity, equity, and inclusion as a core part of our culture.
By applying for this position, you acknowledge and agree that Horizontal Talent may contact you regarding your application using automated technology, including phone calls, SMS/text messages, or email, which may be delivered by our virtual AI recruiter, Alex.
Horizontal is committed to taking affirmative action to employ and advance in employment qualified individuals with disabilities and protected veterans. If you are an individual with a disability and require a reasonable accommodation to complete any part of the application process or participate in the interview process, click here to request accommodation assistance.
All applicants applying must be legally authorized to work in the country of employment.