Multiple choice teaches guessing
Locus has no multiple-choice math. Every numeric or symbolic answer is typed, and a computer algebra system decides whether it is right. The one exception is a word answer, where a short list of options stands in because there is no algebra to check.
Teachers ask why. Options are cheap to grade and impossible to get wrong at scale. The answer is that options change what the student practices. The research on this is old and it points one way.
Retrieval beats review, and the format of the retrieval matters
Start with the reason practice works at all. Roediger and Karpicke (2006) had students read short prose passages, then either restudy them or take a recall test with no feedback. Five minutes later the restudy group was ahead, 81 percent to 75 percent. Two days later that flipped, 68 percent for the tested group against 54 percent. After a week it was 56 percent against 42 percent. In their second experiment, students who studied once and tested three times recalled 61 percent after a week. Students who studied four times and never tested recalled 40 percent.
Two things in that study matter for a practice site. Retrieval outperforms review on any timescale a teacher cares about. And the students who only restudied were the most confident that they would remember. Confidence tracked the losing strategy.
So practice should mean retrieval. The next question is what kind. Kang, McDermott and Roediger (2007) ran the comparison directly. Students read journal papers, then took a multiple-choice quiz, a short-answer quiz, or read a list of statements covering the same facts. In Experiment 1 no feedback was given, and the multiple-choice quiz won. That is not surprising. Students scored 85 percent on the multiple-choice quiz and 56 percent on the short-answer quiz, so the multiple-choice group had seen far more correct answers.
Experiment 2 added corrective feedback to the quizzes, which equalizes exposure to the right answer. The result reversed. The short-answer quiz produced the best retention three days later, beating both the multiple-choice quiz and the read-the-statements condition. Their reading is that harder retrieval pays more, and the harder retrieval is the one with no options on the page.
That is the case for typed answers on a site that grades instantly. Feedback is the condition under which constructed response wins, and software gives feedback in under a second.
Lures are instruction, and they teach the wrong thing
The distractors are not inert. Roediger and Marsh (2005) gave 24 students passages and a multiple-choice test with two, four, or six options per question, then a later cued-recall test where they were warned in strong terms not to guess. Prior testing helped a lot, 46 percent correct against 28 percent for untested facts. But the benefit shrank as the option count grew: 51 percent after two-option questions, 45 percent after four, 43 percent after six. Wrong answers went the other way and rose with option count. Of the errors students produced on the final test, 75 percent were options they had picked on the multiple-choice test. Reading a plausible wrong answer, choosing it, and never being corrected installs it.
Feedback repairs most of the damage. Butler and Roediger (2008) ran the multiple-choice test with immediate feedback, delayed feedback, or none. On a later cued-recall test, correct answers went from 31 percent with no feedback to 45 percent with immediate feedback and 56 percent with delayed feedback. Intrusions of old lures fell from 24 percent to 15 and 14 percent, back down to the level of students who had never seen the lures. Kang and colleagues saw the same repair: a wrong multiple-choice pick was repeated on the final multiple-choice test 65 percent of the time with no feedback and 30 percent with feedback.
The honest summary is that multiple choice with immediate feedback is a decent tool. Multiple choice without it is a machine for teaching wrong facts to the students who are already behind. Most homework platforms sit somewhere between those, and I would rather not depend on which.
There is a math-specific version of this, and it is sharper. Fennell and Foster (2020) gave 902 incoming university students a basic math assessment and randomized the format. The multiple-choice version raised the average score by about one point. Their question 10 shows the mechanism. Asked for a number of people, the most common typed answer was 50, which was the correct percentage. The multiple-choice version did not list 50, so students who had misread the question saw that no option matched, went back, and re-read. The option list caught the error for them. On a real problem it will not.
That is the skill multiple choice trains. Produce something, scan the list, adjust until one matches. It is a good test-taking strategy. It is not mathematics, and it disappears the moment the list does.
What the grader has to do instead
Free-form answers are only worth asking for if you can grade them. The core check is small. Take the student expression and the reference answer, subtract one from the other, simplify the difference, and test whether the result is zero. If it is, the two are the same value written two ways.
That single rule handles the cases teachers hand-check today. A student solving 2x^2 - 3x + 1 = 0 and reporting a root of 1/2 is right. So is 0.5. So is 2/4. Subtract any of them from any other and simplify, and you get zero. None of them need to be listed in advance, and a student who writes -(-1/2) is not punished for it.
Equivalence alone is not enough, because sometimes the form is the answer. Ask a student to factor x^2 - 9 and the answer is (x+3)(x-3). Subtract x^2 - 9 from (x+3)(x-3) and you get zero, so pure equivalence accepts x^2 - 9 as a response to “factor x^2 - 9”. That is the question restated, not solved. So the question carries the form it asked for, and the grader checks structure as well as value: a product of linear factors, not an expanded polynomial. Same for “leave your answer in radical form” and “give the exact value”. The algebra system decides equivalence. The problem decides which shapes count.
Locus grades expressions, equations, intervals, sets, matrices, and ordered lists this way, with the required form enforced per problem. The build cost is real. Every problem has to declare what form it wants, and every answer type needs its own equivalence rule.
What you get back for that is the thing options destroy. A wrong typed answer is a specific wrong answer. It tells you the student dropped a sign, or inverted the fraction, or solved for the wrong variable. A wrong option tells you the student picked C.
References
Butler, A. C., and Roediger, H. L. (2008). Feedback enhances the positive effects and reduces the negative effects of multiple-choice testing. Memory and Cognition, 36(3), 604-616. https://gwern.net/doc/psychology/spaced-repetition/2008-butler.pdf
Fennell, M. A., and Foster, I. R. (2020). Test format and calculator use in the testing of basic math skills for principles of economics: experimental evidence. Institute for International Economic Policy Working Paper 2020-20, George Washington University. https://www2.gwu.edu/~iiep/assets/docs/papers/2020WP/FosterIIEP2020-20.pdf
Kang, S. H. K., McDermott, K. B., and Roediger, H. L. (2007). Test format and corrective feedback modify the effect of testing on long-term retention. European Journal of Cognitive Psychology, 19(4-5), 528-558. https://gwern.net/doc/psychology/spaced-repetition/2007-kang.pdf
Roediger, H. L., and Karpicke, J. D. (2006). Test-enhanced learning: taking memory tests improves long-term retention. Psychological Science, 17(3), 249-255. https://colinallen.dnsalias.org/Readings/2006_Roediger_Karpicke_PsychSci.pdf
Roediger, H. L., and Marsh, E. J. (2005). The positive and negative consequences of multiple-choice testing. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(5), 1155-1159. http://web.archive.org/web/20260429200004/http://psychnet.wustl.edu/memory/wp-content/uploads/2018/04/Roediger-Marsh-2005_JEPLMC.pdf