[quote="rikforto" pid='90786' dateline='1788619086']This account leaves out that the "butchering freedom", as we're calling it, over-generates potential Voynichese cribs. The Rooster example has 17 cribs in the SPS and 10 in Shennong's Classic. [/quote] The longest entry of the SBJ, the "Red Rooster", has 8 (not 10) occurrences of 主. Excluding the fields that apparently were omitted by the Author in all entries, it is 72 hanzi long: 丹雄鸡主崩中漏下赤白沃补虚温中止血通神杀毒辟不祥头主杀鬼肪主耳聋肠主遗溺肶胵裹黄皮主泄利屎白主消渴伤寒寒热翮羽主下血闭鸡子主除热火疮痫痉可作虎魄 The longest single-star parag of the SPS is f105v.32. It has 361 EVA letters in my transcription. Note that 361/72 = ~5, a ratio that seems to be approximately the same for all recipes. Ignoring spaces, there are 15 occurrences of daiin and its allowed variations, marked by brackets below: poarkeeo[daiin]qoairaracphheyqoeedeodyqo[kaiin]qote[dair]aporairapylsheodytairoteeyoteeoolotaiinokeeyqo[kaiin]oraiiraldalsheeo[daiin]chsdqokeeey[dair]o[kaiin]otaiinche[daiin]olkallkl[dain]doeeokcheeoltaiinotcheedychoraiino[daiin]chedyotaiinalkaishd[laiin]sheodokeeodyqoaiinytaiinotairchdaldy[daim]ch[daiin]ockhhyysheyckhysheoqoeeol[kaiin]chsokoltchdysheeeyo[kaiin]araildycheodyoaiirainokshey That's 5 [daiin], 5 [kaiin], 2 [dair], 1 [daim], 1 [dain], 1 [laiin]. I explained before why I allow these but not [taiin], [raiin], [chaiin] etc. Briefly, m is probably an abbreviation of iin or maybe in, and in sloppy cursive handwriting iin may look like ir, and d may look like an l or k (but not a t or other glyphs). Besides those 15, there are 13 additional occurrences of *aiin, *ain, *aim and *air that I do not consider acceptable variations of daiin. [quote]Selecting 10 items from 17 where the order of the selection doesn't matter (because we are going to keep them in the order they appear in the SPS) is 17C10 = 19,448 potential matches.[/quote] Imagine that you have a row of N boxes, and C of them have a coin inside,while the others are empty. You pick M of those boxes at random. You find that F of those M boxes have a coin inside. There are choose(M,F) subsets of F boxes out of the M. That is the number of ways one can get that outcome -- F with coins, M-F empty. If (say) M = 15 and F = 6, choose(M,F) is 11115. That is how may ways there are to find 8 coins in the 15 boxes taken. Does it mean that this outcome -- eight boxes with coins -- is unremarkable? Of course not. That number alone does not mean anything. To tell whether that outcome is significant or not, one must compute its probability taking N into account. The probability that picking M of the N boxes at random will yield exactly F of C coins is choose(C,F)*choose(N-C,M-F)/choose(N,M). How is this relevant to the "Rooster" matching? The Chinese text of the recipe can be viewed as a set of N = 72 boxes, C = 8 of which have a coin (are 主) and the rest are empty (are not 主). A daiin-like string in the candidate paragraph that is centered at index J can be seen as picking the box number round(J/s) of the recipe string, where s is the EVA to hanzi ratio 361/72 = ~5. If parag f105v.32 had nothing to do with the Rooster recipe, the M = 15 occurrences of daiin-like strings in it would be picking a RANDOM subset of the N = 72 boxes. The above formula would then apply: the probability that F = 8 of those M = 15 boxes wouls have a 主-coin would be P = choose(8,8)*choose(64,0)/choose(72,15) = ~0.00000054, or less than one chance in 1'800'000. Now, this is an over-simplified and not quite valid analysis. For one thing, I don't compare the absolute positions of daiin-like strings and 主 characters, but the lengths of the gaps between successive 主 characters and the gaps between the assigned daiin-like strings, scaled by the 1:5 ratio. This complicates the analysis and may change the numbers -- but not enough to change the conclusion. Mreover, when evaluating a match I allow for some variation in the position of the daiin-strings (actually, in the lengths of the gaps). For the Rooster recipe, the lengths of the 9 gaps before, between, and after the 8 assigned daiin-like strings are ??? USE version WITH qo and ir. 8(-7) 110(+1) 11(-4) 15(+0) 46(+9) 20(+0) 45(+3) 25(+0) 52(+0) For example, between the first and the second 主 of the Rooster recipe there are 21 hanzi. By the proportion the 15 daiin-like strings of that paragraph are not all at the precise positions predicted by 1:5 hanzi:EVA ratio. They are off by are choose(15,8) = 15!/8!/7! subsets of 8 ou from the 15 coins. Therefore the probability of that outcome, if the tosses are random, For the second toss, the probability of landing in a different mug is 7 x P; for the next toss, . Therefore, the probability that 8 of the 15 tosses will (0.0375)^8 = ~0.0000000000039, or about 1 chance in 25 billion. But suppose that each coin fell into a different mug. What is the probability of that? The first coin may fall into any mug, so the probability of success in that toss is 15 x 100/40'000. For the second toss the probability of success depends on which mug is 14 x 100/40'000, because the coin must fall into one of the other 14 cups. So the probability of success on all 8 tosses is 15x14x13x12x19x10x9x8 x P^8 = (15!/7!) x P^8 = ~0.0000000000000000000015 or less than 1 chance in 1'400'000'000'000 Now suppose that the mugs are numbered 1 to 15, and each coin fell into a different mug, but in order of increasing mug number. What is the probability of that? Well, for each set of 8 mugs that they could have fallen into, there are 8! = 40'320 possible sequence of those 8 mug numbers, all equally likely, and only one of them is inceasing. Thus the probability is the same as before, divided by 40'320: ~0.000000000000000017, or less than one chance in 58'000'000'000'000'000. (Note that 15!/7!/8! is chose(15,8).) Applying that analogy to the SPS=SBJ claim, the 8 "coins" are the 8 主 in the Rooster entry. The "mugs" are the positions on the SPS parag that correspond to the positions of those 主s in the recipe, calculated assuming an average ratio of ~5 EVA letters for each hanzi. For example, the Voynichese translation of the first 主, which is at position 3.5 in the Chinese text, should be centered about position 5 x 3.5 = 17.5 of the matching SPS paragraph. The second 主 is at position 25.5, so its translation should be at position 5 x 25.5 = 127.5 of the SPS parag. And so on. More precisely, each "mug" is a short interval of positions on the parag centered on the predicted position. Let's say that each interval is plus or minus 5 EVA letter slots on either side. So I count the first 主 as a match if there is a daiin or one of its accepted alternatives centered between positions 12.5 and 22.5 of the EVA string. For the second 主, there must be a daiin-like string centered between positions 122.5 and 132.5. And so on. Since the size of each "mug" is 10 EVA slots and the size of the "table" -- the length of parag f105v.32 -- is 361, the probability P of a given "toss" falling into some "mug" is 10/361 = ~0.027 = ~2.5%. The probability of all 8 tosses landing on mugs is (Actually I consider the distances between consecutive positions rather than the absolute positions. That makes the math a bit more complicated and may change the numbers a bit but not the final conclusion.) The situation is somewhat different from the coin toss above because the keywords are matched sequentially. Here, the probability of the first "coin" landing in any specific "mug" is 10/(372-10) = ~0.0276. But if the first coin falls into the third mug -- that is, 主 is matched to the third candidate on the parag, So what is the probability that the 8 主 "coins" fell into 8 of those 15 ~daiin "mugs" by mere chance? As in the analogy above, it all depends on the "area" of each "mug" and of the "table". Because the output of this function is very sensitive to inputs, this must be tracked carefully. A smaller recipe that generates 10 SPS cribs for 7 hanzi has "only" 120 choices, but if there are two fewer hanzi cribs it exceeds double that. This smaller finding is offset somewhat by the fact that middling recipes have more potential matches; you give 21 for a particular recipe, and the sum of them may approach a large fraction of 19,448. If the size is 2,520 (= 21*(10C7)) or 19,448 is going to have to be pinned down to make a precise evaluation of the odds, but the fact that the "butchering freedom" generates these large search spaces and then you search for one that minimizes the errors has huge consequences for claims about how likely it is that the errors here are small! I did take a look at the problem a few weeks ago, and there is a tendency for your Voynich cribs to significantly exceed the number of Chinese ones, both on average and between proportionally similar entries. I might do this in more detail when you publish a clearer picture of what texts you're using, how you got them, and the parameters of your algorithm. But I feel confident saying the algorithm generates a large match space and then selects a favorable result for your claims. If I'm following your write-ups correctly, you are also running the matching algorithm repeatedly with different crib sets. Transparently, I am not completely clear on the how and why, so it is less clear to me how much this is skewing things. But I can say the decision to selectively exclude cribs is another analytical choice that makes the search space larger than you may realize. If you exclude some cribs and calculate the odds based on that exclusion, you are not calculating the odds for everything you searched. The fact that you have to make the exclusion is important information! There are two points here. The first is that these are analytical choices. They are of a different quality from the last-mile arbitrary interpretations some solvers do, and I do find that sincerely laudable. They are, nonetheless, choices you as the analyst are making. They reflect a pair of texts that you have created based on judgements you have made. Second, the way those choices distort the probability of the match is absolutely insensitive to the strength of your arguments for making these analytical choices. (I am purposely not disputing those choices in this post, which is not to say I think they are all bulletproof.) The fact that you can discard thousands if not 10s of thousands of similarly defined matches must be accounted for in your argument of the odds even if---perhaps especially if!---those choices "make sense". You are correct that scientists can carefully process data to make a match, but that work requires extremely careful guardrails to make sure they are not effectively creating the data. That a number of your choices are invisible to you when evaluating the strength of the match surely means you are underestimating the degree to which you are analyzing changes you have made to the two texts, not the texts themselves! [/quote]