
			<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
			<channel>
			<atom:link href="https://www.thecaliberhunt.com/" rel="self" type="application/rss+xml" />
			<title>Vacancy</title>
			<link>https://www.thecaliberhunt.com</link>
			<description>Vacancy</description>
			<lastBuildDate>Tue, 25 Aug 2026 05:00:55 +0530</lastBuildDate>
			<language>en-us</language>
			<generator>https://www.thecaliberhunt.com</generator>
				<item>
				<title>Data Engineer</title>
				<link>https://www.thecaliberhunt.com/job-openings-for-data-engineer-mumbai-pune-1086668.htm</link>
				<guid>https://www.thecaliberhunt.com/job-openings-for-data-engineer-mumbai-pune-1086668.htm</guid>
				<pubDate>Fri, 26 Aug 2022 00:00:00 +0530</pubDate>
				<description>Technologies / Skills: 
Advanced SQL, Python and associated libraries like Pandas, Numpy etc., Pyspark , Shell scripting, Data- Modelling, Big data, Hadoop, Hive, ETL pipelines and IaC tools like Terraform etc.

Responsibilities:
â€¢	Efficient communication skills to coordinate with users, technical teams and Data\Solution architects.
â€¢	Document technical design documents for given requirements or JIRA stories.
â€¢	Communicate results and business impacts of insight initiatives to key stakeholders to collaboratively solve business problems.
â€¢	Working closely with the overall Enterprise Data &amp; Analytics Architect and Engineering practice leads to ensure adherence with the best practices and design principles.
â€¢	Assures quality, security and compliance requirements are met for supported area.
â€¢	Develop fault-tolerance data pipelines running on cluster
â€¢	Ability to come up with scalable and modular solutions


Required Qualification:
â€¢	1-8 yrs of hands-on experience developing data pipelines for Data Ingestion or transformation using Python (PySpark) /Spark SQL in AWS cloud
â€¢	Experience in development of data pipelines and processing of data at scale using technologies like EMR, Lambda, Glue, Athena, Redshift, Step Functions.
â€¢	Advanced experience in writing and optimizing efficient SQL queries with Python and Hive handling Large Data Sets in Big-Data Environments
â€¢	Experience in debugging, tunning and optimizing PySpark data pipelines
â€¢	Should have implemented concepts and have good knowledge of Pyspark data frames, joins, partitioning, parallelism etc.
â€¢	Understanding of Spark UI, Event Timelines, DAG, Spark config parameters, in order to tune the long running data pipelines.
â€¢	Experience working in Agile implementations
â€¢	Experience with Git and CI/CD pipelines to deploy cloud applications
â€¢	Good knowledge of designing Hive tables with partitioning for performance

Thanks and Regards
HR TEAM</description>
				</item>
			</channel>
			</rss>