We are building a Data Solution Architect, AWS role to shape cloud-native data pipelines for analytics and GenAI enablement on AWS. You will own architecture across ingestion, transformation, storage, and search/retrieval using AWS services and guide teams on delivery standards. Apply to help design and operationalize this data platform end-to-end.
Responsibilities
-
Design scalable data pipelines with AWS Glue for ETL orchestration, AWS Lambda for event-driven compute, and Amazon S3 for data lake storage
-
Define data architecture patterns aligned to the AWS Well-Architected Framework with a focus on reliability, performance, and cost optimization
-
Create technical specifications and architecture diagrams for data platform components
-
Lead implementation of Python and PySpark solutions for large-scale processing and transformation
-
Architect Amazon OpenSearch and vector database solutions that enable semantic search, retrieval-augmented generation, and AI/ML workloads
-
Design data models and pipelines that support advanced analytics, including regression analysis and NLP applications
-
Provide architectural guidance on GenAI strategy, advisory work, and operational integration with existing data infrastructure
-
Design infrastructure patterns that enable machine learning model training, inference, and deployment workflows
-
Advise teams on data preparation and feature engineering approaches for ML/AI use cases
-
Establish CI/CD patterns for data pipeline deployment and infrastructure-as-code practices
-
Collaborate with engineering teams to ensure data platform components meet operational excellence standards
-
Support knowledge transfer and produce technical documentation to ensure sustained delivery
Requirements
-
Solid background with 8+ years of experience in data analytics engineering
-
Hands-on experience writing production-grade PySpark
-
Advanced expertise with AWS Glue, AWS Lambda, and Amazon S3
-
Proven track record using Amazon OpenSearch in real-world solutions
-
English proficiency at B2 level (Upper-Intermediate) or higher
Nice to have
-
Working knowledge of machine learning concepts and workflows
-
Familiarity with CI/CD practices for data and platform delivery
-
Understanding of vector databases and common usage patterns
We offer
-
International projects with top brands
-
Work with global teams of highly skilled, diverse peers
-
Healthcare benefits
-
Employee financial programs
-
Paid time off and sick leave
-
Upskilling, reskilling and certification courses
-
Unlimited access to the LinkedIn Learning library and 22,000+ courses
-
Global career opportunities
-
Volunteer and community involvement opportunities
-
EPAM Employee Groups
-
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn