Computer Vision & Machine Learning Engineer

Dhaka, Bangladesh

Employment Type: Full-time


Job Summary

We are seeking a skilled Computer Vision & Machine Learning Engineer to design, build, and deploy end-to-end multimodal systems for real-time dynamic gesture and continuous sign language recognition and translation. Collaborating closely with our research team, you will bridge state-of-the-art visual feature extraction (3D CNNs, ViTs, pose estimation) with continuous sequence translation to convert live video feeds into accurate, grammatical text. In addition to visual modeling, you will actively architect and optimize cutting-edge NLP pipelines and sequence-to-sequence architectures to handle complex linguistic translation.

Key Responsibilities

  • Model Design & Development: Build and fine-tune Vision Transformers (ViTs), 3D CNNs, Pose Estimation frameworks (e.g., MediaPipe Holistic, OpenPose), and sequence models (Transformers, LSTMs) for dynamic gesture and continuous sign recognition.  

  • Multi-Modal Feature Extraction: Integrate tracking for hands, facial landmarks, and torso pose to capture non-manual linguistic cues alongside manual signs.  

  • Data Engineering & Augmentation: Design robust pre-processing pipelines, handle domain shift (lighting, skin tones, camera angles, occlusion), and apply spatial-temporal data augmentation strategies.  

  • Sequence & Translation Modeling: Bridge visual feature extraction with Natural Language Processing (NLP) models to map sign glosses and continuous video sequences to grammatical target text.

  • Optimization & Deployment: Quantize and optimize models (e.g., ONNX, TensorRT, CoreML, TFLite) for low-latency, real-time inferencing on edge devices, web, and mobile environments.  

  • Evaluation & Benchmarking: Establish robust metrics (BLEU, ROUGE, WER, Frame-Level Accuracy) for visual recognition and sign translation performance.

  • Deployment & API Integration: Participate in the deployment of machine learning models into production environments using containerization technologies (e.g., Docker) and create RESTful/gRPC APIs (e.g., FastAPI, Triton Inference Server) for seamless integration with other systems and applications.

  • Ethics & Inclusivity: Collaborate with Deaf community advisors and sign language experts to ensure datasets and models reflect real-world linguistic diversity and signing variations.

  • Research & Innovation: Stay up-to-date with the latest advancements in machine learning and AI technologies. Contribute to research initiatives, publish papers, and participate in conferences (e.g., CVPR, ICCV, ECCV, NeurIPS) as appropriate.

Required Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, Information Technology, or a related field.

Required Skills

  • Bachelor’s or master's degree in Computer Science.

  • 2–5 years of hands-on experience in model training and machine learning engineering

  • Strong problem-solving skills and the ability to learn quickly in a rapidly evolving research domain.

  • Solid understanding of machine learning algorithms and techniques (e.g., deep learning, 3D CNNs, CTC, Vision Transformers, pose estimation, sequence modeling, etc.).

  • Strong programming skills in Python and proficiency with libraries such as PyTorch, TensorFlow, OpenCV, Transformers, and scikit-learn.

  • Proven experience in developing and deploying machine learning models and systems across domains such as computer vision, NLP, and deep learning.

  • Hands-on experience with containerization technologies (e.g., Docker) and creating APIs (FastAPI, Triton Inference Server) for system integration.

  • Experience with building scalable data pipelines, spatial-temporal data augmentation, and working with large-scale video/multimodal datasets.

  • Strong communication skills and ability to collaborate effectively in a cross-functional team environment.

Preferred / Additional Skills

  • In-depth experience with Triton Inference Server for multi-model orchestration, along with custom model quantization, pruning, and low-latency serving on edge/mobile environments.

  • Solid background in building end-to-end MLOps pipelines (e.g., CI/CD for machine learning, automated model retraining, tracking via MLflow/WDB, and container orchestration).

  • Deep familiarity with low-level CUDA setup, GPU environment configuration, parallel computing libraries (e.g., TensorRT, cuDNN), and kernel optimization for high-throughput video processing.

  • Experience leveraging distributed training frameworks (e.g., PyTorch DDP, DeepSpeed) to process massive, high-resolution video streams and multimodal datasets efficiently.

  • Proven track record of publishing research papers in top-tier Computer Vision or NLP conferences.

Key Competencies

  • Analytical & Technical problem solving.

  • Ability to work independently and take ownership of tasks.

  • Good communication and documentation skills.

  • Domain adaptation & robustness.

  • Cross-Functional collaboration & empathy.

Job Location

  • The primary work location is Dhaka, Bangladesh.

Salary Information

  • An attractive remuneration package will be offered for deserving candidates.

  • The base salary will be determined based on the candidate's years of experience (DOE).    

Compensation & Other benefits

  • Flexible leave/vacation policy.

  • Public holidays as declared by Bangladesh Government.

  • Lunch Facilities: Partially Subsidized.

  • Salary Review: Yearly.

  • Festival Bonus: 2 (Yearly).


To Apply: Please submit your resume and a cover letter detailing your experience and why you’re a good fit for the role.