Data Scientist
Must-have:PythonGitAWSAzureGoogle CloudCloudBackendDataAI
We are seeking a hands-on Data Scientist with strong experience in building and deploying AI/ML-powered applications. This role goes beyond research: you'll work closely with cross-functional teams to develop, integrate, and scale intelligent features in real-world products. You will also handle administrative and operational tasks that keep the data and AI team running smoothly.
Responsibilities:
- Design, develop, and deploy AI/ML models to power core features of our applications.
- Collaborate with Product, Engineering, and Design teams to translate business requirements into technical solutions.
- Build end-to-end pipelines from data ingestion and preprocessing to model inference and integration.
- Implement and optimize computer vision models (e.g., YOLO) for real-time applications.
- Integrate LLMs using frameworks like Langchain and manage embeddings via vector databases (e.g., FAISS, Chroma).
- Develop and implement Retrieval-Augmented Generation (RAG) systems using LLMs and vector databases to enhance contextual relevance.
- Develop APIs and backend services for AI features using Python frameworks like Flask, FastAPI, or Django.
- Perform feature engineering, data cleaning, and model evaluation to ensure high performance in production.
- Create clear and actionable dashboards, reports, or tools for monitoring and troubleshooting ML models.
- Stay current with best practices in AI engineering, deployment, and monitoring.
- Manage team administration and operations, including project documentation, meeting schedules and minutes, task tracking, and progress reporting.
- Maintain organized records of project files, datasets, model documentation, and reports for easy access and audit.
- Handle operational administration such as tracking cloud and software subscriptions, licenses, and budgets, and preparing related reports or purchase requests.
- Coordinate with other departments and external vendors on schedules, requirements, and follow-ups.
- Prepare regular operational reports (project status, resource usage, team KPIs) for management.
Qualifications:
- Bachelor's or Master's degree in Data Science, Computer Science, Statistics, Physics, or related fields.
- 2+ years of hands-on experience building and deploying ML-powered applications in production.
- Strong proficiency in Python, including ML libraries (scikit-learn, TensorFlow, PyTorch).
- Solid understanding of machine learning tasks: regression, classification, clustering, time series, etc.
- Familiar with LLM integration tools such as Langchain, and handling embeddings with FAISS or Chroma.
- Experience designing or implementing RAG pipelines that combine LLMs with retrieval systems.
- Solid experience building REST APIs and AI services with Flask, FastAPI, or Django.
- Proficient in working with large-scale data using SQL and NoSQL databases.
- Understanding of model deployment, versioning, and monitoring in production environments.
- Comfortable using cloud services (AWS, GCP, Azure) for model training and deployment.
- Familiar with tools for reproducibility and collaboration (Git, MLflow, DVC, etc.).
- Experience with Computer Vision, especially using YOLO (v5/v7/v8) or similar object detection frameworks.
- Strong organizational and time-management skills, with attention to detail in documentation and record-keeping.
- Proficiency with office and collaboration tools (Google Workspace/Microsoft Office, Jira/Trello/Notion, etc.).
- Good written and verbal communication skills for coordinating across teams and vendors.
Skills: SQL, Machine Learning, Data Analytics, Python, Data analysis