Join PrimeGrid
Returning Participants
Community
Leader Boards
Results
Other
drummers-lowrise
|
Message boards :
Number crunching :
Tremendous number of proof tasks
| Author |
Message |
|
|
|
My understanding from the Primegrid main page is that all non-sieving applications except for AP-27 use fast-proof tasks rather than a full double-check. It also appears, looking at the task results, that each main task gets assigned a single proof task. Thus, my expectation is that overall there is (or should be) about a 1:1 ratio between main task & proof tasks, and thus we should expect to receive each in (about) equal numbers. However, I am noticing what seems to be a tremendously higher number of proof tasks than main tasks. To quantify that, looking through the 1000 tasks my devices have pulled down most recently, I am seeing 159 main tasks (& thus, 841 proof tasks). Now, I do understand that the proof tasks are a "eat your vegetables" situation, because PG won't work without the double-checking process, but this seems somewhat excessive, and is somewhat discouraging because the proofs don't (I think?) have any shot of finding primes. Any idea why I am seeing such a skew and how I might bring it closer to balanced? Again, I'm plenty happy to do my "fair share" of these tasks....but when I'm getting 85% proof tasks, that doesn't sound quite like a fair share :)
(posting here since this seems to be the case for both GFN & PPS(E), but if there's somewhere better to ask, feel free to move this) | |
|
mackerel Volunteer tester
 Send message
Joined: 2 Oct 08 Posts: 2982 ID: 29980 Credit: 791,367,422 RAC: 39,992
                                       
|
|
Had a look on one of my systems out of interest. I counted 86/200 main tasks, so much closer to half.
I wonder if there is something more than luck going on. For example, could the frequency of server contact or cache sizes influence what you get? | |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
|
It’s mostly luck. Maybe a tiny bit of timing.
The DC tasks always have a higher priority, so if any exist when you request work, you’ll get them first. But they are very short, and the same rules apply to everyone, so it will even out.
____________
My lucky number is 75898524288+1 | |
|
|
|
|
I been watching that on my daily driver. Sometimes I go nearly an hour without a single proof task, and sometimes I end up with a dozen or more in a row. It's all good since, as Michael said, they're quick. | |
|
|
|
|
Now that I've had a chance to SSH into a couple of my systems to look at a bit more data, I do think I was overreacting above.
Only 10% on this one:
user@host:~$ boinccmd --get_state | grep "WU name" | grep c | wc -l
502
user@host:~$ boinccmd --get_state | grep "WU name" | wc -l
5024
And about 20% here:
user@host:~$ boinccmd --get_state | grep "WU name" | wc -l
5004
user@host:~$ boinccmd --get_state | grep "WU name" | grep c | wc -l
1090
Related, does PG have any way to bulk export data through the website (an API endpoint, or a way to download a CSV/JSON file of results)? I'm the sort that would greatly like the ability to do some analysis on the data, and that'd be much easier with bulk access to it. | |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
Related, does PG have any way to bulk export data through the website (an API endpoint, or a way to download a CSV/JSON file of results)? I'm the sort that would greatly like the ability to do some analysis on the data, and that'd be much easier with bulk access to it.
Not for results, no. There’s a handful of pages that will take xml as an argument and output xml instead of html, and while that’s slightly easier to parse than html, it’s not really that much different. And the results page doesn’t do that anyway. (Don’t ask me which pages do this; I don’t know and it’s not documented as far as I know.)
Scraping the pages is possible; we don’t mind provided you do it single threaded, one page at a time. No launching multiple page requests at once.
FWIW, nearly all candidates in fast-DC projects have exactly one main task and one DC task, so in the long run you simply have to get very close to a fifty fifty split. It’s unlikely there’s any way, intentional or otherwise, to influence which kind of tasks you will get. If there’s DC tasks available, you always get them first, but they’re much shorter, and they’re equal in number to the main tasks, so despite what you might think you see, your computers are only spending about 1% of their computing time on the DC tasks.
____________
My lucky number is 75898524288+1 | |
|
RafaelVolunteer tester
 Send message
Joined: 22 Oct 14 Posts: 1002 ID: 370496 Credit: 1,083,511,878 RAC: 180,601
                                  
|
Not for results, no. There’s a handful of pages that will take xml as an argument and output xml instead of html, and while that’s slightly easier to parse than html, it’s not really that much different. And the results page doesn’t do that anyway. (Don’t ask me which pages do this; I don’t know and it’s not documented as far as I know.)
Scraping the pages is possible; we don’t mind provided you do it single threaded, one page at a time. No launching multiple page requests at once.
FWIW, nearly all candidates in fast-DC projects have exactly one main task and one DC task, so in the long run you simply have to get very close to a fifty fifty split. It’s unlikely there’s any way, intentional or otherwise, to influence which kind of tasks you will get. If there’s DC tasks available, you always get them first, but they’re much shorter, and they’re equal in number to the main tasks, so despite what you might think you see, your computers are only spending about 1% of their computing time on the DC tasks.
I find manually requesting a big cache, letting it run dry, then repeating will get me a lot more of main tasks. When requesting only one or two tasks because you're running on 0, chances are you will get that one DC task from someone that just reported a result; if you request for 50 tasks at once, sure, you might get a couple hanging, but then you're also getting a bunch of non DC the server is assigning to you. | |
|
|
|
It’s mostly luck. Maybe a tiny bit of timing.
The DC tasks always have a higher priority, so if any exist when you request work, you’ll get them first. But they are very short, and the same rules apply to everyone, so it will even out.
The same rules don't apply to everyone. People who do not have any type of cache are almost certainly going to end up with fewer double check tasks relative to the number of main tasks in aggregate since they are getting new tasks one at a time instead of multiple at a time.
For me, I've found that the issue is the worst for smaller tasks since more tasks are fetched simultaneously for them.
I would be interested for the admins to calculate the main:double check ratio for every user across all projects. My guess is that for larger projects it will be closer to a normal distribution, but for GFN 16 and PPSE it will appear to be a slight bimodal distribution.
I requested adding the number of double check tasks performed to the table on the user page a long time ago. James proposed putting it on https://www.primegrid.com/_recents.php, but that never came to fruition.
____________
| |
|
|
|
I would be interested for the admins to calculate the main:double check ratio for every user across all projects. My guess is that for larger projects it will be closer to a normal distribution, but for GFN 16 and PPSE it will appear to be a slight bimodal distribution.
I requested adding the number of double check tasks performed to the table on the user page a long time ago. James proposed putting it on https://www.primegrid.com/_recents.php, but that never came to fruition.
That would be some good data to have, for me at least. Shame it never happened. | |
|
|
|
Related, does PG have any way to bulk export data through the website (an API endpoint, or a way to download a CSV/JSON file of results)? I'm the sort that would greatly like the ability to do some analysis on the data, and that'd be much easier with bulk access to it.
Not for results, no. There’s a handful of pages that will take xml as an argument and output xml instead of html, and while that’s slightly easier to parse than html, it’s not really that much different. And the results page doesn’t do that anyway. (Don’t ask me which pages do this; I don’t know and it’s not documented as far as I know.)
Scraping the pages is possible; we don’t mind provided you do it single threaded, one page at a time. No launching multiple page requests at once.
FWIW, nearly all candidates in fast-DC projects have exactly one main task and one DC task, so in the long run you simply have to get very close to a fifty fifty split. It’s unlikely there’s any way, intentional or otherwise, to influence which kind of tasks you will get. If there’s DC tasks available, you always get them first, but they’re much shorter, and they’re equal in number to the main tasks, so despite what you might think you see, your computers are only spending about 1% of their computing time on the DC tasks.
Cheers. I'll look into writing something that will let me parse the data & toss it into a database so I can run some stats / do some analysis on it from there.
Not for results, no. There’s a handful of pages that will take xml as an argument and output xml instead of html, and while that’s slightly easier to parse than html, it’s not really that much different. And the results page doesn’t do that anyway. (Don’t ask me which pages do this; I don’t know and it’s not documented as far as I know.)
Scraping the pages is possible; we don’t mind provided you do it single threaded, one page at a time. No launching multiple page requests at once.
FWIW, nearly all candidates in fast-DC projects have exactly one main task and one DC task, so in the long run you simply have to get very close to a fifty fifty split. It’s unlikely there’s any way, intentional or otherwise, to influence which kind of tasks you will get. If there’s DC tasks available, you always get them first, but they’re much shorter, and they’re equal in number to the main tasks, so despite what you might think you see, your computers are only spending about 1% of their computing time on the DC tasks.
I find manually requesting a big cache, letting it run dry, then repeating will get me a lot more of main tasks. When requesting only one or two tasks because you're running on 0, chances are you will get that one DC task from someone that just reported a result; if you request for 50 tasks at once, sure, you might get a couple hanging, but then you're also getting a bunch of non DC the server is assigning to you.
Interesting. Sounds like way too much trouble :-)
I would be interested for the admins to calculate the main:double check ratio for every user across all projects. My guess is that for larger projects it will be closer to a normal distribution, but for GFN 16 and PPSE it will appear to be a slight bimodal distribution.
I requested adding the number of double check tasks performed to the table on the user page a long time ago. James proposed putting it on https://www.primegrid.com/_recents.php, but that never came to fruition.
That would be some good data to have, for me at least. Shame it never happened.
+1 -- sounds like that would be nice | |
|
|
|
|
I agree, I'm getting over 90% of proof tasks vs main tasks as per my rough calculation
pretty annoying to get no chance at finding a prime whatsoever
my cache settings are 0 and 0
no cache whatsoever forced by Boinc | |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
I agree, I'm getting over 90% of proof tasks vs main tasks as per my rough calculation
pretty annoying to get no chance at finding a prime whatsoever
my cache settings are 0 and 0
no cache whatsoever forced by Boinc
Even if 90% DC tasks was sustainable (it's not) you'd still be spending 90% of your CPU time on the main tasks, since they are usually 64 or 128 times longer than the DC tasks.
So you still have a 90% chance of working on a task that CAN find a prime, at any given instant.
Once you get some better luck, it will be closer to 98% or 99%.
____________
My lucky number is 75898524288+1 | |
|
RafaelVolunteer tester
 Send message
Joined: 22 Oct 14 Posts: 1002 ID: 370496 Credit: 1,083,511,878 RAC: 180,601
                                  
|
I agree, I'm getting over 90% of proof tasks vs main tasks as per my rough calculation
pretty annoying to get no chance at finding a prime whatsoever
my cache settings are 0 and 0
no cache whatsoever forced by Boinc
To be fair, you're making your life a lot harder by selecting multiple subprojects. Remember, the server assigns top priority to double check tasks, so before sending you main tasks, it will send all DCs for GFN 16, then 17, then 18, etc. before even having a chance at sending something else. Trim it down to a single subproject per CPU / GPU and you'll see your ratio get closer to 50/50. | |
|
|
|
|
-- | |
|
|
|
To be fair, you're making your life a lot harder by selecting multiple subprojects. Remember, the server assigns top priority to double check tasks, so before sending you main tasks, it will send all DCs for GFN 16, then 17, then 18, etc. before even having a chance at sending something else. Trim it down to a single subproject per CPU / GPU and you'll see your ratio get closer to 50/50.
that's the funniest and less logical answer I ever heard to any question ;-)
there's no "remember" if it is not known in the first place, LMAO
Even if 90% DC tasks was sustainable (it's not) you'd still be spending 90% of your CPU time on the main tasks, since they are usually 64 or 128 times longer than the DC tasks.
So you still have a 90% chance of working on a task that CAN find a prime, at any given instant.
This, on the other hand, makes sense.
Thank you | |
|
|
|
|
I think it is fine for the server to give priority to DC tasks when only a single task is requested, but if the server is sending back multiple tasks, I think it should try to give them in a 50/50 ratio.
____________
| |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
I think it is fine for the server to give priority to DC tasks when only a single task is requested, but if the server is sending back multiple tasks, I think it should try to give them in a 50/50 ratio.
50/50 by count?
Or 50/50 by time?
1 main task and 50 DC tasks is still better than 50/50 by time.
(It isn't really possible to change how it works.)
____________
My lucky number is 75898524288+1 | |
|
|
|
I think it is fine for the server to give priority to DC tasks when only a single task is requested, but if the server is sending back multiple tasks, I think it should try to give them in a 50/50 ratio.
50/50 by count?
Or 50/50 by time?
1 main task and 50 DC tasks is still better than 50/50 by time.
(It isn't really possible to change how it works.)
50/50 by count
It's still not equitable when everyone should be doing first check and proof checks in equal ratios. At this point I'm probably just going to put together a script to abort proof checks outside of the 50/50 ratio and run it on a cron schedule.
____________
| |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
|
How about some hard data?
select userid, u.name, count(*) total, sum(r.name not like '%c\_%') main, sum(r.name like '%c\_%') dc, concat(format(sum(r.name like '%c\_%')/count(*)*100,2),'%') 'dc %' from result r join user u on r.userid=u.id where r.userid in (1654901, 29980, 1342315, 53948, 35074, 392025, 370496) group by r.userid;
+---------+-----------------------------+--------+--------+-------+--------+
| userid | name | total | main | dc | dc % |
+---------+-----------------------------+--------+--------+-------+--------+
| 29980 | mackerel | 4979 | 1804 | 3175 | 63.77% |
| 35074 | Michael Millerick | 64425 | 27840 | 36585 | 56.79% |
| 53948 | Michael Goetz | 5388 | 2734 | 2654 | 49.26% |
| 370496 | Rafael | 2361 | 1261 | 1100 | 46.59% |
| 392025 | [AF>EDLS]zOU | 12899 | 2958 | 9941 | 77.07% |
| 1342315 | Holdolin | 67089 | 27616 | 39473 | 58.84% |
| 1654901 | Aperture_Science_Innovators | 201499 | 116382 | 85117 | 42.24% |
+---------+-----------------------------+--------+--------+-------+--------+
That's all the people who have posted in this thread. Note that the thread starter has received more main tasks than DC tasks. In fact, among this group, he's got the best ratio of main to DC tasks.
The other Michael, who is presumably busy writing his script, is indeed getting more DC tasks, albeit by a modest 57%.
I've got the most even distribution, and [AF>EDLS]zOU has the worst, with 3/4 of his tasks being DC.
But again, recall that the DC tasks are very, very short, so all of us are almost exclusively actually working on main tasks. That is, after all, the whole point of the fast DC system.
____________
My lucky number is 75898524288+1 | |
|
|
|
|
Thank you Michael, I appreciate the time you took and the very reasonable/sensible answer you gave me ;)
Crunch on :D | |
|
|
|
|
I don't see the issue. So there are folks out there that think they should not be running DC tasks that potentially help someone else find a prime? Last I looked, this whole project is geared to finding primes. If running lots of DC tasks helps that, and it does, then that is goodness. | |
|
|
|
|
To be honest I would totaly run an old celeron or pentium chromebox or a low power free cloud tier instance on just DC tasks only to help you guys out if it was possible!
Much better than it spending 10-20k seconds on a single ppse task | |
|
|
|
How about some hard data?
Can the hard data be added to https://www.primegrid.com/_recents.php or some other more visible page so that none of us need to conjecture about it again in the future? My ratio certainly doesn't feel as close to 50% as it apparently is from casually watching BOINC download tasks on my second screen.
____________
| |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
How about some hard data?
Can the hard data be added to https://www.primegrid.com/_recents.php or some other more visible page so that none of us need to conjecture about it again in the future? My ratio certainly doesn't feel as close to 50% as it apparently is from casually watching BOINC download tasks on my second screen.
tl;dr: No. If we do that, half the people will be needlessly upset at any given moment. As things are right now, most people get it that over the long term you're going to get 50/50.
Only if you want to join the staff and take on the job, in perpetuity, of answering the continuous stream of complaints from people receiving 50.01% or more DC tasks.
There's some information that doesn't actually help anyone, but causes some people to become obsessive of certain details, that we decide not to make public. This is not for technical reasons, but rather because it aggravates people and brings no benefit. Note that this is not the only such information in this category. Back before fast DC, we could have, and I wanted to, provide detailed information about your wingmen's progress. Specifically, you could have seen the trickle information from their tasks. This, however, created nothing but additional stress, so that information was not displayed. The DC ratio falls into the same category.
I realize this sounds like I'm treating you like children. Perhaps I am, after all, TdP is essentially a big game, no? But I'm really treating you like humans -- very passionate, competitive adults. I have access to more information, but I'm also a participant, just like you, and I definitely do not want to have this information at my fingertips. I can be happy in the certainty that over the large enough data set I *must* be getting roughly a 50/50 split. In small data sets, however, that will vary. I don't want to see the short term numbers, since half the time those will annoy me. It's better to look at the big picture, which has to be 50/50.
____________
My lucky number is 75898524288+1 | |
|
|
|
|
Hear, hear
____________
Thanks,
Jim
| |
|
|
|
It's better to look at the big picture, which has to be 50/50.
I disagree that you can categorically say it has to be 50/50 over the long term. I'm thinking about my computer's behavior on GFN16. Whenever it fetches work, it fetches at most one main task, however the double check tasks are so short that it fetches multiple of those. I'm quite sure that if you break out my ratios per sub project you are going to see GFN16 skew much more towards double check tasks that the other subprojects do.
E.g. here is the latest fetch, which pulled 39 double check tasks and no main tasks.
This is almost certainly solved by not running smaller projects, and I am going to avoid them once TdP is done. It also may be solved by setting resource share to zero and only allow a single task to be queued up at once--but I am disappointed that the project refuses to give us access to the information that would allow users to see whether or not tweaking settings allows users to get meaningfully closer to the expected ratio between main and proof checks.
____________
| |
|
|
|
|
Some people have too much time on their hands! | |
|
Conan   Send message
Joined: 24 Mar 09 Posts: 1246 ID: 37336 Credit: 319,182,500 RAC: 223,893
                         
|
|
For any one that appear to be getting far more DC tasks than Main tasks, please remember that there is a challenge on and the amount of tasks being downloaded , especially the shortest tasks, is huge.
So the amount of processed Main tasks is also huge, so a Lot of DC tasks have to be sent out to verify those Main tasks.
if you are a bit unlucky you might get a lot of DC tasks some of the time and you can get quite a few Main tasks at others.
It evens out but taking a local spot check can skew the overall 50/50 one way or the other.
The DC tasks have to be done and someone has to do them, or no Prime will be proven.
Just enjoy the challenge at hand you can still snag a Mega even with lots of DC tasks.
They are very short so don't get hung up about them.
Have a bit of fun.
I am still here even though I have yet to get a Prime or Mega in any TdP I have been in over the years, I still enjoy it.
Conan
____________
| |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
This is almost certainly solved by not running smaller projects, and I am going to avoid them once TdP is done. It also may be solved by setting resource share to zero and only allow a single task to be queued up at once--but I am disappointed that the project refuses to give us access to the information that would allow users to see whether or not tweaking settings allows users to get meaningfully closer to the expected ratio between main and proof checks.
It is a fallacy to believe you can affect whether you get main or DC tasks in a meaningful way.
____________
My lucky number is 75898524288+1 | |
|
|
|
This is almost certainly solved by not running smaller projects, and I am going to avoid them once TdP is done. It also may be solved by setting resource share to zero and only allow a single task to be queued up at once--but I am disappointed that the project refuses to give us access to the information that would allow users to see whether or not tweaking settings allows users to get meaningfully closer to the expected ratio between main and proof checks.
It is a fallacy to believe you can affect whether you get main or DC tasks in a meaningful way.
We don’t have access to the data that would allow that to be demonstrated in a meaningful way—unless you want us to forcefully scrape it one page of 20 tasks at a time from the web server.
All that us commoners know is that the server forcefully prioritizes proof tasks over main tasks, which leads to the trend I screenshotted above, where for small projects the server will send you dozens of proof tasks at once but would only send a single main task due to the difference in expected duration.
____________
| |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
All that us commoners know is that the server forcefully prioritizes proof tasks over main tasks, which leads to the trend I screenshotted above, where for small projects the server will send you dozens of proof tasks at once but would only send a single main task due to the difference in expected duration.
You're not thinking it through. Where do the DC tasks come from?
____________
My lucky number is 75898524288+1 | |
|
|
|
All that us commoners know is that the server forcefully prioritizes proof tasks over main tasks, which leads to the trend I screenshotted above, where for small projects the server will send you dozens of proof tasks at once but would only send a single main task due to the difference in expected duration.
You're not thinking it through. Where do the DC tasks come from?
You're trying to redirect the conversation to the project level, the concern here is at the user level. Yes, there has to be a 50/50 ratio at the project level based on how the work is generated. There ought to be a 50/50 ratio for every user in order to be equitable. My thesis is that the way the server sends out work doesn't allow that 50/50 ratio to be achieved for small projects that have very fast proof tasks unless a user goes out of their way to modify their settings to be similar to what we had to do before in order to maximize the likelihood of getting a first--reduce resource share to 0 and effectively not allow any work units to be stored up.
As an example, I scraped all of my completed GFN16 tasks out of the web server. I had 3666 main tasks and 7312 proof tasks, which happens because I am able to fetch far more of the proof tasks at once because of their short expected four second duration. I'll keep running this script nightly, put the results in a spreadsheet, and post the information back to demonstrate that the ratio can be manipulated based on user settings.
____________
| |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
|
It's just luck. Changing your cache settings has no effect on which tasks you get. Over a large enough period of time, you'll get close to a 50/50 ratio.
I'm not going to argue about this anymore. If you don't want to believe me, don't.
____________
My lucky number is 75898524288+1 | |
|
|
|
To be fair, you're making your life a lot harder by selecting multiple subprojects. Remember, the server assigns top priority to double check tasks, so before sending you main tasks, it will send all DCs for GFN 16, then 17, then 18, etc. before even having a chance at sending something else. Trim it down to a single subproject per CPU / GPU and you'll see your ratio get closer to 50/50.
that's the funniest and less logical answer I ever heard to any question ;-)
there's no "remember" if it is not known in the first place, LMAO
Not known? He probably said "remember" because Michael Goetz already mentioned that the server gives top priority to DC tasks earlier in this very same thread.
____________
1 GFN-18, 1 PPSE, 5 SGS, and 5 GFN-15 primes | |
|
|
|
I don't see the issue. So there are folks out there that think they should not be running DC tasks that potentially help someone else find a prime? Last I looked, this whole project is geared to finding primes. If running lots of DC tasks helps that, and it does, then that is goodness.
There's no "fairness" in finding primes. trying to get a 50/50 ratio (number not time) increases the chances of finding one
I've not found a single one.
This user (https://www.primegrid.com/show_user.php?userid=317224) #1 in my team with 1,082,504,591.18 credit hasn't found a single one either in 10 years.
This user ( https://www.primegrid.com/show_user.php?userid=1671125) with 2 weeks of membership and 201,019.94 credit has found one.
For each main task, there is a DC task.
DC tasks are MUCH faster than main tasks.
so as there's the same amount of main and DC tasks, there's no reason for an uneven sending of tasks from the server.
Unless some users are "deliberately" storing and holding main tasks vs DC tasks. and the "aborted" DC tasks have to be resent (again and again) until they get allocated to a user who is going to accept them.
And in the meantime, a large bunch of main taks is getting stored with some users and the matching DC tasks gets allocated to other users ?
I may have missed something.
Not known? He probably said "remember" because Michael Goetz already mentioned that the server gives top priority to DC tasks earlier in this very same thread.
yes, this week.
Was it documented anywhere before ?
Is it documented anywhere at all ?
That's what I mean; I can't know something that's been told 2 days prior when we're discussing that's been going on for much longer than 2 days :D | |
|
mackerel Volunteer tester
 Send message
Joined: 2 Oct 08 Posts: 2982 ID: 29980 Credit: 791,367,422 RAC: 39,992
                                       
|
|
IMO this isn't really worth worrying about.
Earlier it was stated double check units are either 1/64 or 1/128 the length of a main task. Let's take 1/64 as the worst case.
If we take a ratio of 1 main task and 1 double check to be "ideal", you're spending 98.5% of the time doing main tasks. If you never do a double check unit, you gain 1.5% main tasks done relative to that state.
Keeping 1 main task to 1 DC task as reference 100% throughput, how much does that change if you vary the number of DC units per main?
2 DCs per main is a 1.5% reduction in main unit throughput.
3 DCs per main is a 3.0% reduction in main unit throughput.
4 DCs per main is a 4.4% reduction in main unit throughput.
Going the other way:
1 DC per 2 mains is a 0.8% increase in main unit throughput.
If DCs are 1/128 instead of 1/64 then you can probably halve the numbers above. | |
|
|
|
IMO this isn't really worth worrying about.
If we take a ratio of 1 main task and 1 double check to be "ideal", you're spending 98.5% of the time doing main tasks. If you never do a double check unit, you gain 1.5% main tasks done relative to that state.
Ultimately I agree with you.
But, I've had up to 77% of my tasks being DC and 23% being main.
So I wonder where are the main tasks going or who's filtering out the DC tasks from their received WU ?
Because, ultimately, it "should" be even in numbers
unless the fastest GPU/CPU are getting the main tasks and everyone else the DC to accelerate main tasks completion ?
After all, the fastest CPU/GPU are 10/100 time faster than some other... so it would make sense to give them the main tasks.
Which is great for the project, but highly unfair for the user in term of reward :D | |
|
|
|
IMO this isn't really worth worrying about.
Earlier it was stated double check units are either 1/64 or 1/128 the length of a main task. Let's take 1/64 as the worst case.
If we take a ratio of 1 main task and 1 double check to be "ideal", you're spending 98.5% of the time doing main tasks. If you never do a double check unit, you gain 1.5% main tasks done relative to that state.
Keeping 1 main task to 1 DC task as reference 100% throughput, how much does that change if you vary the number of DC units per main?
2 DCs per main is a 1.5% reduction in main unit throughput.
3 DCs per main is a 3.0% reduction in main unit throughput.
4 DCs per main is a 4.4% reduction in main unit throughput.
Going the other way:
1 DC per 2 mains is a 0.8% increase in main unit throughput.
If DCs are 1/128 instead of 1/64 then you can probably halve the numbers above.
DCs on smaller subprojects are 1/16 or 1/32 of the duration, which changes your numbers substantially. A 6% difference in throughput is absolutely worth worrying about—we worried about far less when finding a prime was a race, now it is a different kind of race with different mechanisms to optimize the outcome of the race. Instead of racing your wingman, it is a race against the black box prioritization of proof tasks by the server.
____________
| |
|
|
|
|
You're not thinking it through. Where do the DC tasks come from?[/quote]
...LOL Michael...as always done by you:
YOU made my day! Thanks!!
____________
NUMBER CRUNCHING... will not save us
| |
|
Honza Volunteer moderator Volunteer tester Project scientist Send message
Joined: 15 Aug 05 Posts: 2082 ID: 352 Credit: 9,805,570,281 RAC: 3,090,130
                                                   
|
This user (https://www.primegrid.com/show_user.php?userid=317224) #1 in my team with 1,082,504,591.18 credit hasn't found a single one either in 10 years.
This user ( https://www.primegrid.com/show_user.php?userid=1671125) with 2 weeks of membership and 201,019.94 credit has found one.
I may have missed something.
You sure did missed something and frankly, quite a lot....
user 317224 (damotbe) may have a bit over 1082 milions credit, which is in this case a mostly irrelevant when comparing to finding primes.
About 50% of total credit was done on PPS Sieve, where you can't find any primes in principle.
Another 42% was done WW, where basically no primes were found - by anybody.
So, PPS Sieve and WW are doing bunch of credit on GPUs but doesn't find a prime.
This leave us only with about 8% for project that may find a primes.
From those, GFN has a bit oiver 5% of total credit.
Digging deeper, those tasks were GFN Extreme (DYFL), which again give a nice credit bonus but not a single prime in whole history of DYFL by anyone. Even SoB would have a better chance.
btw, a 875. position for GFN is not much credit at all or about 0.4% of what 1. position did.
If domatbe is trying to find a prime, this is completely bogus and disastrous strategy.
Ratio between main and DC task is completely irrelevant in this case, it's basically bad strategy chosing those subprojects is one's goal is to find a prime.
User 1671125 just got lucky finding a smallest posible GFN prime.
And now trying to find GFN17.
Well, it takes dozens of thousand candidates to find one on average.
Doing about 20 GFN17 a day on single AMD 7950X CPU, I guess it would take about 1500 days on average to find one.
Not to discourage but we will see.
Ratio between main and DC task is also irrelevant there, just got lucky.
____________
My stats | |
|
|
|
To be honest I would totaly run an old celeron or pentium chromebox or a low power free cloud tier instance on just DC tasks only to help you guys out if it was possible!
Much better than it spending 10-20k seconds on a single ppse task
I have an AMD Phenom that I could add to that effort. It takes 2-3 minutes to run a PPSE double check and over an hour for a main task. | |
|
|
|
You sure did missed something and frankly, quite a lot....
Thank you :)
I appreciate the time you took, I'm not that familiar with PG sub-project to have had a relevant look at the details, but I'm sure very glad you did.
I very much appreciate it, and understand now. | |
|
|
|
But, I've had up to 77% of my tasks being DC and 23% being main.
So I wonder where are the main tasks going or who's filtering out the DC tasks from their received WU ?
I wonder if this computer is affecting the numbers at all, by having over 30,000 tasks in progress at any one time, the vast majority of which are never completed...
It's certainly keeping some of my credit hostage, as DCs for some of my bigger tasks have gone to this computer and then become stuck there for days on end, only to end up timing out.
____________
1 GFN-18, 1 PPSE, 5 SGS, and 5 GFN-15 primes | |
|
mackerel Volunteer tester
 Send message
Joined: 2 Oct 08 Posts: 2982 ID: 29980 Credit: 791,367,422 RAC: 39,992
                                       
|
DCs on smaller subprojects are 1/16 or 1/32 of the duration, which changes your numbers substantially.
I didn't check for any specific project at the time and went by what was posted earlier. Looking at one of my systems on PPSE it does look like DC in those cases may be 1/32. Roughly, it is 2x what the 1/64 numbers I gave earlier were, and are:
0% DC: No DCs is a 3% throughput advantage.
33% DC: 1 DC per 2 mains is a 1.5% increase in main unit throughput.
50% DC: 1 DC per 1 main reference condition.
67% DC: 2 DCs per main is a 2.9% reduction in main unit throughput.
75% DC: 3 DCs per main is a 5.7% reduction in main unit throughput.
80% DC: 4 DCs per main is a 8.3% reduction in main unit throughput.
Personally I'm still not bothered by this. | |
|
Vato Volunteer tester
 Send message
Joined: 2 Feb 08 Posts: 997 ID: 18447 Credit: 2,403,191,125 RAC: 111,129
                                      
|
|
compared to the old double check system, i prefer all of those options to a 66% first ratio
____________
| |
|
WezHSend message
Joined: 9 Jun 11 Posts: 220 ID: 101605 Credit: 2,045,642,137 RAC: 4,185
                                       
|
compared to the old double check system, i prefer all of those options to a 66% first ratio
I second that. Now my poor old dual core Pentium G640 have change to find prime, and it did! https://www.primegrid.com/workunit.php?wuid=1058451997
Before fast dc first ratio was about zero all time... | |
|
|
|
compared to the old double check system, i prefer all of those options to a 66% first ratio
I second that. Now my poor old dual core Pentium G640 have change to find prime, and it did! https://www.primegrid.com/workunit.php?wuid=1058451997
Before fast dc first ratio was about zero all time...
Hey, Gratz on the prime. I'd imagine that is exactly why the new system was put in place. Heck, personally I don't mind doing proof tasks at all. Nothing wrong with helping out the other cruncher waiting on their validation =] | |
|
Michael Goetz Volunteer moderator Project administrator
 Send message
Joined: 21 Jan 10 Posts: 14712 ID: 53948 Credit: 1,051,161,470 RAC: 387,049
                                           
|
compared to the old double check system, i prefer all of those options to a 66% first ratio
I second that. Now my poor old dual core Pentium G640 have change to find prime, and it did! https://www.primegrid.com/workunit.php?wuid=1058451997
Before fast dc first ratio was about zero all time...
Hey, Gratz on the prime. I'd imagine that is exactly why the new system was put in place. Heck, personally I don't mind doing proof tasks at all. Nothing wrong with helping out the other cruncher waiting on their validation =]
That’s half the reason. The other part is that effectively it makes searching for primes nearly twice as fast since half of our available computing power isn’t being spent on double checking.
____________
My lucky number is 75898524288+1 | |
|
|
|
|
Doesn't bother me, we are all in the same boat. | |
|
Davina   Send message
Joined: 13 Feb 12 Posts: 3494 ID: 130544 Credit: 2,937,568,616 RAC: 438,447
                                   
|
|
Short tasks? Enjoy them. | |
|
|
|
|
Could not agree more!
____________
Thanks,
Jim
| |
|
|
|
|
Sorry if my question is repetition. I have been away for quite some time. Precisely since the changes in the double-checking occured. And I have read all messages posted here, but I think there is no clear answer to my question.
As far as I understand, the double-checking has not been eliminated, it simply takes less time now to double-check a prime. So it means that there is someone finding a prime and someone else checking it, as before. Besides checking primes involving less resources, what else changed? Does the double-checker continues to get recognition? I mean, if my computer verifies a prime found by someone else, will that prime I checked be listed on my list of double-checked primes? This is my main question, as I dont mind doing the short WU, what I dont like is to see my list of double-checked primes grow. In fact, personally, I would prefer the list of double-checked primes to disappear entirely. I prefer not knowing which and how many primes I double-checked. So does that list continue to show the new double-checked primes or not?
Thanks
____________
| |
|
|
|
|
You are not recognized as a double checker, because you are just doing some quick math stuff and not searching for a prime. Double check tasks are purged quicker than regular tasks, so your list will not going to grow. This also helps keep the database smaller and hopefully more manageable.
____________
Lucky numbers 121*2^4553899-1 and 3756801695685*2^666669±1
My 12 min movie https://vimeo.com/manage/videos/502242 | |
|
|
|
|
I'll add to Pooh Bear's answer:
With a fast DC task, you aren't actually verifying a prime, just that the work done was done correctly. The PG server will do a verification run of its own, followed by the top5k site (if eligible for submission).
On the one hand, no more secondhand glory; on the other hand, you are (very nearly*) always the prime finder regardless of your system's speed of computation; and on the gripping hand, almost double the prime finding power of our systems!
*There is the very rare occasion that the original task assignee has some sort of error that is reported to the server so that a second task is generated and sent out. But then, that first task is recovered from the error, and so one of the tasks will necessarily be the second received. It's an admin question, but to my knowledge, none of those tasks have actually found a prime, but I suppose if they did, returner #2 would get the prime on their double checker list.
____________
Eating more cheese on Thursdays. | |
|
|
|
|
Thanks to both of you for the clarification.
____________
| |
|
Message boards :
Number crunching :
Tremendous number of proof tasks |