Summary
What you’ll impact
Our organization is seeking a senior AI/Software Engineer in New York City to develop a large‑scale document intelligence platform using Python, OCR, LLMs, and GPU infrastructure. The role involves building production services, data pipelines, and hybrid on‑prem/cloud solutions, with a part‑time commitment of about 20 hours per week and potential for more hours based on performance.
Responsibilities
What you'll do
- Build production Python services, APIs, and data-processing pipelines
- Process large volumes of PDFs, HTML, images, scans, and structured/unstructured data
- Build OCR, extraction, parsing, normalization, and document-classification pipelines
- Extract tables, entities, relationships, metadata, citations, and structured records
- Build ETL and high-volume batch-processing systems
- Develop hybrid search, embeddings, vector search, RAG, reranking, and evidence retrieval
- Integrate Claude, Qwen, OpenAI, Gemini, and open-source models
- Deploy and serve local LLMs on NVIDIA H100/H200 GPUs
- Work with vLLM, PyTorch, Hugging Face, CUDA, quantization, batching, and GPU inference optimization
- Build and maintain the AWS/cloud deployment
- Work with PostgreSQL, OpenSearch, S3, Redis, queues, and related data infrastructure
- Build APIs, webhooks, Stripe integrations, authentication, and third-party integrations
- Write automated tests covering OCR, extraction, retrieval, data pipelines, and AI outputs
- Debug complex AI, data, backend, and infrastructure issues
- Document architecture and implementation decisions
Requirements
What you’ll bring
- Education: If you did your undergrad in India, save your time and DO NOT APPLY
- You will have to submit a W-9 when hired, no OPT, W-8, or other types of sponsorship
- Work Authorization: Candidates must be legally authorized to work in the United States under an arrangement compatible with a W-9 contractor engagement. NextGen Coding Company cannot provide OPT or employment/visa sponsorship for the role.
- Location: New York City — Hybrid / In-Person Required
- Strong Python engineering experience
- Strong backend and data-engineering fundamentals
- Experience building production software
- OCR / document intelligence experience
- Experience processing PDFs, images, and large datasets
- Production experience with LLMs, RAG, embeddings, and vector search
- Experience deploying open-source models on NVIDIA GPU infrastructure
- Understanding of H100/H200-class inference environments
- Strong AWS experience
- PostgreSQL / SQL
- Docker and Linux
- REST APIs and third-party integrations
- Ability to independently own difficult engineering problems
- Applicants should provide: Resume or LinkedIn
- Applicants should provide: GitHub and/or technical portfolio
- Applicants should provide: Examples of production AI/data systems built
- Applicants should provide: Experience with Python, OCR, LLMs, AWS, and GPU infrastructure
- Applicants should provide: Specific NVIDIA GPU/model deployment experience, if applicable