It works when the same class of decision recurs often enough to build a track record, when people's strengths and weaknesses are already known and openly discussed, and when the decision is consequential enough to justify the overhead. Dalio's own version at Bridgewater ran on years of accumulated data about what each person was like.
It fails when nobody has a real track record, which is most early-stage startup decisions. With no history, believability collapses into seniority or confidence, and you have rebuilt autocracy with extra steps. It also fails when the scoring is done in public without trust already in place, where it reads as ranking people rather than ranking opinions. And Dalio is candid that the cultural cost is real: he says the transparency at Bridgewater was widely perceived from outside as a cult.
Where operators split. Dalio's whole apparatus assumes decisions improve when you systematize the judgment and surface it as data. Jason Fried argues close to the opposite for product calls, saying people are fundamentally feeling creatures and that the feel of a thing often outweighs the spreadsheet. Both are describing real conditions. Dalio is talking about repeatable, high-stakes, evidence-rich decisions. Fried is talking about taste, where the track record you need does not exist yet.