MIT Builds AI Tool for Novice Coders in Military Settings

An experiment shows that using AI chatbots, a cadet with no coding experience built a battlefield application in three months.

By Central
Highlights
  • Cadet Joshua Lynch with no prior coding experience created the ROMAD-AI application using only AI chatbot prompts.
  • The experiment revealed key challenges like loss of hierarchical focus, requiring strategic prompting skills.
  • The research suggests vibe-coding is best for rapid prototyping, not production deployment, and needs security review.

The U.S. Department of the Air Force and MIT Lincoln Laboratory have demonstrated that novice coders with no prior programming experience can build functional software applications using generative AI chatbots, a finding with significant implications for how military personnel and other nontechnical specialists can rapidly prototype tools tailored to their operational needs. The experiment, conducted through the MIT AI Accelerator’s Phantom Program, tested whether a complete beginner could rely entirely on prompting AI language models—a process known as “vibe-coding”—to move from concept to working application without formal training in software development.

What Is Vibe-Coding and Why It Matters for Military Software Development

Vibe-coding describes a workflow in which a user writes virtually no code manually but instead guides a generative AI chatbot through natural-language prompts to produce, debug, and refine software. The approach bypasses the traditional, resource-intensive military software pipeline, which often requires formal requirements documentation, specialized developers, and lengthy procurement cycles. For service members who understand tactical problems intimately but lack technical training, the ability to create custom tools directly could accelerate innovation at the unit level and reduce dependence on centralized development teams.

The Experiment: A Cadet With Zero Coding Experience Builds a Battlefield Application

U.S. Air Force cadet Joshua Lynch, who had no background in software development, set out to determine whether he could create a fully functional application for his tactical team. With mentorship from Laura Niss, a technical staff member in MIT Lincoln Laboratory’s Embedded and AI Systems Group, Lynch spent three months using three commercial AI chatbots: Anthropic’s Claude, OpenAI’s ChatGPT, and Google’s Gemini. He operated primarily through the chatbots’ web-based chat interfaces rather than integrated development environments, relying entirely on prompts to generate, test, and modify code.

The original ambition was ambitious: an application called the Remote Operating Modular Augmentation Device, or ROMAD-AI, that would offer AI-assisted target recognition, modular intelligence and surveillance capabilities, autonomous striking coordination, and battlefield communication management. Over the course of the project, Lynch completed several professional development courses in AI and studied both military and civilian applications of the technology to inform his prompting strategy.

Key Technical Challenges and Workarounds Learned Through Experience

Lynch encountered several recurring limitations that required him to develop specific prompting strategies. The chatbots frequently lost hierarchical focus, modifying unrelated sections of code when a change was requested for a specific function. He learned to break problems into smaller, discrete components, frame questions with precise context, and steer conversations back on track when the model strayed from the objective. These techniques consumed most of the three-month timeline but ultimately allowed him to produce a working prototype.

The final version of ROMAD-AI was built using Google AI Studio, which integrates directly with the Gemini API and provides an AI-aware development environment. Due to time constraints and the inherent limitations of the language models, the scope of the application was reduced from a real-time battlefield assistant to a document-processing tool capable of analyzing tactical maps and generating mission-planning documents through a VLM-powered chatbot interface. The prototype was not secure enough for the intended operational environment, but it demonstrated the feasibility of the approach.

How Perception of AI Changed Over Time

One of the project’s most revealing findings was how Lynch’s perception of the AI systems evolved. After starting with an ambitious vision shaped by media portrayals of AI capabilities, he gained a pragmatic understanding of what current models can and cannot reliably do. His expectations were significantly recalibrated by the end of the project. Niss tracked changes in his perceptions across multiple dimensions, including likeability, anthropomorphism, and perceived intelligence, noting that Claude showed more stability across these traits during the study period compared to ChatGPT. Lynch reported that AI served as a useful tutor and coding assistant but became noticeably less reliable on topics he knew well, where he could spot inaccuracies directly.

Security Risks and the Persistent Need for Code Review

A critical incident during the project underscored the security risks of deploying AI-generated code without rigorous review. Lynch discovered that the final application was sending input documents to the Gemini AI model for analysis rather than parsing them locally on his machine—a behavior he had not intended and did not initially detect. This type of data-handling oversight is particularly concerning in military contexts where sensitive information is involved. Niss noted that while AI can generate large volumes of functional code, thorough code review remains a bottleneck and a non-negotiable step for any application handling confidential data or critical operations.

Broader Implications for Nontechnical Users in Specialized Fields

The project’s results suggest that AI chatbots are most effective as prototyping assistants rather than full production tools for nontechnical users, especially when sensitive information is involved. Niss observed that these systems can be powerful for helping nontechnical experts communicate problems and potential solutions to technical teams, effectively serving as a translation layer between operational needs and software implementation. “No matter how good AI gets, I think we’ll always need to collaborate to get to the best solutions for the most important problems,” she said, reinforcing the idea that human expertise in both the operational domain and software engineering remains essential.

What This Means for Military and Enterprise AI Adoption

For organizations considering similar approaches, the experiment offers several actionable insights. First, vibe-coding can dramatically lower the barrier to prototyping for personnel who understand specific operational problems but lack coding skills. Second, successful outcomes depend heavily on the user’s ability to learn effective prompting strategies, including problem decomposition and conversational steering. Third, any application intended for use with sensitive data must undergo thorough security review by qualified engineers, and the current generation of AI chatbots cannot be trusted to handle data-handling correctly without explicit validation. Fourth, the scope of what can be achieved with pure prompt-based development is narrower than what an experienced developer can produce, suggesting that the best use case is rapid prototyping rather than production deployment.

The research was sponsored by the Department of the Air Force Artificial Intelligence Accelerator under Cooperative Agreement Number FA8750-19-2-1000, and it demonstrates a concrete path for empowering domain experts across military and civilian sectors to create software tools without waiting for formal development cycles. For any organization with staff who understand the problem deeply but lack programming experience, the takeaway is clear: invest in teaching prompt engineering fundamentals, establish a lightweight code review process, and treat AI-generated prototypes as the starting point for collaboration with technical experts—not as final deliverables.

Share This Article