A weighted scorecard you hand to every AI vendor in an RFP, so the decision that lands on your P&L is made on evidence, not on whoever gave the best demo. Same categories, same weights, scored side by side.
Read this first · Why this exists
A demo is a sales asset, not a due-diligence artifact. Every AI vendor can show a clean dashboard on curated data. What that presentation cannot tell you is the only thing that decides whether the money you are about to commit returns anything: will the tool read your actual floor, run on your actual machines, and survive contact with your actual team. This scorecard forces those questions into the open and grades every vendor against the same rubric.
There are no invented statistics on this page. The one external figure here is a public regulation with a link. The weights in the scorecard are a starting framework you set to match your plant and your risk, not a claim about the market. Everything a vendor scores is measured against evidence they give you, not against a benchmark we made up.
Why a scorecard beats a demo
When AI purchases fail in a plant, they rarely fail because the model was bad. They fail because the tool could not reach the data, could not talk to the machines already on the floor, or asked more of the team than the team could give, and it quietly went unused. None of those failure modes show up in a demo. They show up in month four, after the money is spent and the political capital with it.
A demo optimizes for the vendor's strengths. A scorecard optimizes for your risks. It puts every vendor on the same axes, weighted the way you decide matters, so a strong answer on rollout cannot be hidden behind a flashy interface, and a great interface cannot paper over the fact that the thing will never see your real records. The output is one comparable number per vendor and, more usefully, a paper trail of exactly where each one is weak. That is what turns a gut call into a board-defensible decision.
Score it against the sequence, not the pitch
The scorecard is built on one premise every ready plant learns the hard way: AI runs on data a system can read, and it has to be digitized, connected, and unified before any agent has something to stand on. A vendor who cannot connect to your floor is not selling you AI. They are selling you a screen that needs your team to keep feeding it by hand. The five categories below are ordered so the questions that most often kill a purchase come first.
The five categories, and why each one carries weight
Here is the shape of it. Each category is one axis every vendor is graded on, with a suggested weight you adjust to your own plant. The full weighted scorecard and the underlying RFP question set are below, sent to your work email or revealed here on this page.
Category 01
Data connectivity: can it read your floor at all
Grades: reach into your actual records
Whether the tool can pull from the systems and paper you run today, or whether value depends on your team keying data into it forever. This is the single most predictive category, so it carries the most weight.
Category 02
Machine reach: does it work on any PLC
Grades: brand-agnostic connection to your assets
Whether the tool connects to the mixed-vintage controllers you already own over a standard like OPC UA, or whether it only works on one brand, or needs every line retrofitted before it does anything.
Category 03
Rollout burden: what it costs your team
Grades: the load on your people, not just the invoice
The real price is your operators' and engineers' time. Who does the integration, how long until working software, how much change is forced on the floor, and what happens when the vendor's team goes home.
Category 04
Output that feeds something
Grades: whether the data leaves the tool
Whether the tool's output flows into your ERP, MES, quality and scheduling systems, or dead-ends in one more dashboard nobody outside the tool can use. A number trapped in a screen changes no decision.
Category 05
Honest ROI: math you can audit
Grades: return you can verify on your own inputs
Whether the vendor's ROI rests on your measured numbers and named assumptions you can check, or on borrowed benchmarks and percentages from other plants that do not describe yours.
Want to know whether your own floor is even ready to be scored against these vendors? Start with the AI Readiness Checklist, then put a dollar figure on the gaps with the ROI Calculators & Tools.
The full scorecard
Get the weighted scorecard and RFP question set.
You have seen the five categories. Enter your work email and the full weighted scorecard opens right here on this page, with the exact RFP questions to put to each vendor and the scoring scale to grade them. A copy goes to your inbox to run your evaluation from.
Suggested weights for all five categories, set to change on your own risk
The RFP questions to send each vendor, with what a strong answer looks like
A 0 to 5 scoring scale and how to read the total across vendors
Work email only. We use it to send the scorecard and nothing else you did not ask for. Unsubscribe anytime.
Unlocked. The full scorecard is open below, and a copy is on its way to your inbox. If you checked the box, a Harmony engineer will reach out to pressure-test your shortlist with you.
The weighted scorecard
Use it like this: send every shortlisted vendor the questions in each category, score each category 0 to 5 from the evidence they return, multiply by the weight, and total. The weights below are a starting point. Set them yourself before you see a single vendor, so the rubric is not bent to fit a favorite after the fact. If recalls and audits are your sharpest risk, push data connectivity and output higher. If your floor is a mix of old and new controllers, machine reach earns more.
01 · Data connectivity30%
02 · Machine reach, any PLC20%
03 · Rollout burden on your team20%
04 · Output that feeds something15%
05 · Honest, auditable ROI15%
Suggested total100%
The 0 to 5 scoring scale
0No answer, or the answer proves the category is not solved. A hard stop worth flagging on its own.
1 to 2Answered in principle, but with conditions, roadmap promises, or "we can build that" rather than something that runs today.
3 to 4Answered with a concrete method and at least one reference to a plant like yours where it is already live.
5Answered with evidence you can verify yourself: a live connection, a working artifact, a number built from your own inputs.
Category 01 · Data connectivityWeight 30%
Can it read your floor at all
The question underneath every other question. If the tool cannot reach your real records, nothing downstream matters. Ask, then grade the evidence, not the reassurance.
AskWhich of our existing systems will you connect to directly, and which will require our people to enter data by hand? Name them.Strong answer: a specific list of your ERP, MES, quality and warehouse systems they connect to, and an honest short list of what stays manual.
AskHow do you get records that start as paper on the floor into the tool, batch sheets, quality checks, travelers, downtime logs?Strong answer: a real capture method at the station, not "your team types it in at end of shift."
AskOn day one of the pilot, what percentage of our weekly production data can your tool actually see without new manual entry? How did you arrive at that number?Strong answer: a figure derived from a walk of your floor, with the method shown, not a confident round number.
AskWhen a record only exists as handwriting today, what does your tool do with it?Strong answer: an acknowledgment that it cannot read what is not digitized, and a plan to digitize it, not a claim that it reads paper by magic.
Why this carries the most weight
An agent that cannot read the record cannot act on the record. A copilot answering from a system that holds a third of your week gives confident answers about a third of your plant, and confident wrong answers are more dangerous than no answer. A vendor who scores low here is not selling you AI. They are selling you a screen your team has to keep feeding by hand, forever, and the labor to feed it will erase the return.
Category 02 · Machine reachWeight 20%
Does it work on any PLC
Your floor is a mix of controllers bought over decades from different vendors. Grade whether the tool meets that reality or demands you rebuild it.
AskOur lines run controllers from more than one brand and more than one decade. Which of ours do you connect to today, and how?Strong answer: connection over an open standard such as OPC UA, and specifics on the brands and vintages already in the field.
AskDoes your tool require us to replace or retrofit controllers, add gateways, or standardize on one brand before it works?Strong answer: works on what you own, with any added hardware named, priced, and kept to a minimum.
AskFor a machine that only exposes counts on an HMI or a clicker today, how do you get that data out?Strong answer: a concrete method for freeing data that stops at the panel, not "that machine is not supported."
AskWhat happens to the parts of our floor your tool cannot reach on day one? Do they stay dark, or is there a path?Strong answer: an honest map of what connects now versus later, with the later part sequenced.
What a low score here means
Many controllers can already share data at the PLC over open standards, so the common problem is not a plant that cannot measure. It is a plant whose measurements stop at the panel, and a vendor who only supports one brand leaves most of your floor dark. A tool that forces a rip-and-replace before it produces value has moved a software decision into a capital-equipment decision, and the timeline and cost belong on a different page of your budget entirely.
Category 03 · Rollout burdenWeight 20%
What it actually costs your team
The license is the smallest cost. The real bill is your operators' and engineers' time and the disruption to a floor that has to keep shipping. Grade the load on your people.
AskWho does the integration work, your team or ours? How many hours of our people's time does a rollout take, and in what roles?Strong answer: the vendor does the heavy lifting, with a named, bounded ask of your staff.
AskHow long from signature to working software our team can use on the floor, not to a kickoff deck?Strong answer: a timeline in weeks with a defined first working milestone, not an open-ended program.
AskHow much do our operators have to change how they already work, and what is your plan when the floor resists a new screen?Strong answer: meets the floor where it is, with a real change-management approach, not "they will adapt."
AskWhen your implementation team leaves, what do we own and run ourselves, and what stays dependent on you?Strong answer: you own the system and the data, with dependency on the vendor named and limited.
Why the invoice is the wrong number to watch
Plenty of AI purchases die not from cost but from asking a stretched team for time it does not have, and then quietly going unused. A vendor who cannot answer how many of your people's hours a rollout consumes has not planned the rollout. For reference on what a bounded engagement can look like, a Harmony pilot is a fixed $15,000 to $20,000 one time, 4 to 6 weeks, with working software by the end of the pilot, forward-deployed engineers doing the integration alongside your team. Use that as a yardstick for whether a vendor's answer is concrete or open-ended.
Category 04 · Output that feeds somethingWeight 15%
Does the data leave the tool
A number that changes a decision has to reach the person or system that makes the decision. Grade whether the output flows out or dead-ends in one more dashboard.
AskWhere does your tool's output go? Can it write back into our ERP, MES, quality and scheduling systems, or does it only display?Strong answer: named write-back or export into the systems your team already works in.
AskCan we get our own data out in a standard, sortable electronic format whenever we ask, including for an audit or a recall?Strong answer: yes, on demand, in a portable format, with no lock-in tax.
AskIf we run a one-lot trace, does your tool cut the time to pull records across receiving, production, quality and shipping, or is it a separate silo?Strong answer: it shortens the real cross-system trace, rather than adding a new place to check.
Why this is a compliance question, not just a convenience
Output that cannot leave the tool is output that changes nothing, and it becomes a liability the moment a regulator or a recall is on the clock. Under the FDA Food Traceability Rule, which implements section 204 of the Food Safety Modernization Act, covered firms must make required traceability records available to an authorized FDA representative within 24 hours of a request, and during an outbreak, recall, or other public health threat must provide the required information in an electronic sortable spreadsheet within that same window. A tool that traps your data in its own screen cannot help you meet either condition. Grade export and write-back as if an audit depends on it, because for many manufacturers it does.
Source: 21 CFR 1.1455, paragraphs (c)(1) and (c)(3)(ii), FDA Food Traceability Rule.
Category 05 · Honest ROIWeight 15%
Math you can audit
Every vendor will show a return. Grade whether the return is built from your numbers and stated assumptions you can check, or from borrowed benchmarks that describe someone else's plant.
AskShow me your ROI model. Which inputs are our measured numbers, and which are assumptions you supplied? Name every assumption.Strong answer: a model built on your inputs, with each assumption labeled and defensible, not a slideware payback.
AskWhich claimed savings can we verify from our own records after go-live, and how will we measure them?Strong answer: a short list of savings tied to metrics you already track, with a way to check them.
AskWhat has to be true for this ROI to fail? Where have similar plants not seen the return, and why?Strong answer: an honest account of failure modes, not a claim that it always works.
Treat borrowed percentages as a red flag, not evidence
An average across other people's plants does not tell you what will happen on yours, and a vendor who leads with a headline percentage is asking you to trust a number you cannot audit. The only ROI worth weighing is one built from figures you can pull from your own floor. A vendor who insists on modeling the return on your measured inputs, and who can name the assumptions that would sink it, is more trustworthy than one whose case rests on a benchmark from a plant you have never seen. Price the gaps yourself with the ROI Calculators & Tools and hold every vendor's model against your own.
Where a vendor's answers put you: the three phases
The categories map onto the same three phases every plant moves through. Most AI vendors sell as though your floor is already in Phase 3. The scorecard reveals which phase you are actually in, and a vendor whose value depends on data you do not yet have connected is selling you the last phase before you have built the first.
Phase 1
Lay the Data Foundation · Digitization
Every pen-and-paper record digitized at the station, every software system connected, and all of the data unified into one live layer. The digital transformation starts here.
Phase 2
Production & Operations Scale
Factory operations turn proactive: live sensors and machine data, the AI scheduling board, predictive maintenance before failure.
Phase 3
AI-Native Operations
Agents across the floor and the back office act on the live layer: quality signals, reports, copilots. Humans approve.
Before you score a single vendor, confirm your own floor can support what they are selling. The AI Readiness Checklist is the plain list you work through first, and the Manufacturing Paper Audit counts the gaps station by station.
Want a second set of eyes on your shortlist?
The same scorecard is how our forward-deployed engineers pressure-test a plant's AI options, on-site, Phase 1 first, because that is the order it has to happen in. See what the live layer looks like.