AI / ML engineer focused on model optimization, quantization, and deployment.
I help teams take models from “works in research” to “runs efficiently in production.” My work includes PyTorch to ONNX export/debugging, INT8/INT4 quantization, calibration, accuracy recovery after compression, latency/memory benchmarking, and fine-tuning/evaluation workflows.
Best fit: startups or teams with ML/AI models that are too slow, too expensive to serve, hard to deploy, or losing accuracy after quantization/compression.
Interested in freelance, consulting, part-time, or full-time opportunities around model optimization, edge AI, inference efficiency, and applied ML systems.
Remote: Yes
Willing to relocate: Yes
Technologies: Python, PyTorch, ONNX, ONNX Runtime, quantization, QDQ, INT8/INT4, model compression, calibration, accuracy recovery, computer vision, LLM fine-tuning, LoRA/QLoRA, evaluation pipelines, Docker, Linux
Résumé/CV: Available on request
Email: johnzakkam2592@gmail.com
AI / ML engineer focused on model optimization, quantization, and deployment.
I help teams take models from “works in research” to “runs efficiently in production.” My work includes PyTorch to ONNX export/debugging, INT8/INT4 quantization, calibration, accuracy recovery after compression, latency/memory benchmarking, and fine-tuning/evaluation workflows.
Best fit: startups or teams with ML/AI models that are too slow, too expensive to serve, hard to deploy, or losing accuracy after quantization/compression.
Interested in freelance, consulting, part-time, or full-time opportunities around model optimization, edge AI, inference efficiency, and applied ML systems.