Testing DeepSeek-V4-Flash Official Release with Codex: A 30-Question Hardcore Benchmark
Built a pure-standard-library benchmark harness with Codex, then made real API calls to DeepSeek-V4-Flash (0731 official) at 2026-08-02 10:51 to run 30 self-built questions. Result: 30/30 correct, 59/59 coding test cases passed, 30-question cost under 5 fen, ~3s average latency, 84% reasoning tokens. Includes the official 9-benchmark comparison and a price showdown (V4-Flash output ~1/90 of Opus 4.8). A hands-on benchmark with reproducible, auditable raw data, including limitations and known weaknesses.