Start with AI projects you can measure and reverse

At PayIt, our AI team is testing how agents can support internal teams and take on new work. The recommendations in this post come from that experience, including an effort to upgrade roughly 200 live services to Java 25.
AI leads NASCIO’s 2026 list of state CIO priorities. As government leaders explore where AI can help, we recommend starting with projects that are measurable, reversible, and easy to monitor.
Start with two questions
- If the AI produces the wrong result, how will we find out, and how quickly?
- Can we undo what it did?
Strong answers to both questions give teams room to test, learn, and build confidence.
What that looked like at PayIt
We are using AI agents to help upgrade roughly 200 live services to Java 25. The work includes changing code, deploying across environments, testing, and debugging. Completing it manually would normally occupy a team for several quarters.
This project works well for AI because the results can be evaluated objectively. The code compiles, and passes its tests, or it fails. It is clear when a problem has occurred, making it more straightforward to correct the work and move forward.
That AI use case is different from tasks such as generating summaries, where accuracy may depend on human judgment.
A practical way to group projects
Start with internal projects that are easy to measure, monitor, and reverse, such as software upgrades or reconciliation triage. Resident-facing tools like service search, form prefill, or translation should include human review. AI that influences benefit eligibility, fraud decisions, denials, or penalties requires the greatest scrutiny because the consequences are more significant and harder to reverse.
Allow for the unexpected
We initially tried putting agents on fully predetermined paths. When an agent encountered something its instructions did not cover, it did not always recognize that it should stop or adjust.
Agents need enough flexibility and context to handle unexpected conditions. Connecting them to our internal knowledge base also allowed later work to benefit from lessons recorded during earlier upgrades. That flexibility still requires testing, boundaries, and human oversight.
What residents think about government AI
This risk framework matters even more when AI interacts with residents. In the summer of 2026, we surveyed 1,000+ U.S. adults and found that comfort varied by task. It fell when AI was responsible for decisions about benefit eligibility.
want an AI agent handling payments or applications.
want to know when they are interacting with AI.
People who had used government AI tools rated their experiences positively. Agencies should start with limited uses, clearly label AI interactions, and explain when staff review the work.
Five questions to ask before moving forward
-
1
How will we identify a wrong result?
-
2
Can we reverse it?
-
3
Does the AI assist a person or make the decision?
-
4
How will we explain where AI is involved?
-
5
How will staff use the time the project saves?
Choose an initial project with clear success measures, strong oversight, and manageable consequences if something goes wrong. That creates a practical foundation for expanding AI responsibly.
A final note on resident-facing AI
Resident-facing projects require additional care because their effects may be harder to identify or reverse. Start with contained use cases that reflect residents’ comfort levels, clearly disclose AI’s role, monitor performance, and keep people accountable for the outcome.
You might also like
Useful GovTech insights — right in your inbox
Get articles and insights from our monthly newsletter.




