about the company
This role is with a leading global technology service provider. They specialize in delivering advanced connectivity solutions, digital services, and enterprise technology transformations across the region.
about the job
- Design, build, and optimize scalable batch and streaming data pipelines using Databricks and Apache Kafka.
- Perform data transformation and cleansing using PySpark and SQL to meet technical and business needs.
- Lead the implementation and operationalization of vector databases, knowledge bases, and RAG architectures for enterprise generative AI applications.
- Integrate data from diverse source systems (APIs, databases, files, streaming) and collaborate with cross-functional teams to define ingestion patterns.
- Oversee end-to-end production readiness, including observability, incident triage, root-cause analysis, and operational runbooks.
- Ensure data governance, cloud security, access control, and compliance standards are enforced across all pipelines and knowledge platforms.
skills and experience required
... - Bachelor’s degree in Computer Science, Software Engineering, or a related quantitative field.
- 5 to 8 years of hands-on experience in data engineering, cloud-scale data platforms, or analytics delivery.
- Proficiency in Python, PySpark, and SQL for data manipulation, transformation, and validation.
- Practical experience building and deploying knowledge base/RAG solution stacks and vector embeddings for AI use cases.
- Strong understanding of cloud technologies and orchestration tools (e.g., Databricks, Delta Lake, Microsoft Fabric, CI/CD pipelines).
- Excellent problem-solving, documentation, and communication skills to lead technical discussions with stakeholders.
To apply online please use the 'apply' function, alternatively you may contact Sophia Lim at 92487488.
(EA: 94C3609/ R26163753 )