[SPARK-29641][PYTHON][CORE] Stage Level Sched: Add python api's and tests #28085

tgravescs · 2020-03-31T20:05:29Z

What changes were proposed in this pull request?

As part of the Stage level scheduling features, add the Python api's to set resource profiles.
This also adds the functionality to properly apply the pyspark memory configuration when specified in the ResourceProfile. The pyspark memory configuration is being passed in the task local properties. This was an easy way to get it to the PythonRunner that needs it. I modeled this off how the barrier task scheduling is passing the addresses. As part of this I added in the JavaRDD api's because those are needed by python.

Why are the changes needed?

python api for this feature

Does this PR introduce any user-facing change?

Yes adds the java and python apis for user to specify a ResourceProfile to use stage level scheduling.

How was this patch tested?

unit tests and manually tested on yarn. Tests also run to verify it errors properly on standalone and local mode where its not yet supported.

SparkQA · 2020-03-31T20:10:22Z

Test build #120650 has finished for PR 28085 at commit 2b515c8.

This patch fails Python style tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2020-03-31T21:29:08Z

Test build #120651 has finished for PR 28085 at commit a052427.

This patch fails Python style tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2020-04-01T04:43:42Z

Test build #120653 has finished for PR 28085 at commit 6a90fbe.

This patch fails from timeout after a configured wait of 400m.
This patch merges cleanly.
This patch adds no public classes.

tgravescs · 2020-04-01T13:14:08Z

test this please

SparkQA · 2020-04-01T20:00:06Z

Test build #120677 has finished for PR 28085 at commit 6a90fbe.

This patch fails from timeout after a configured wait of 400m.
This patch merges cleanly.
This patch adds no public classes.

tgravescs · 2020-04-01T20:25:30Z

the tests are timing out running these:
-- Gauges ----------------------------------------------------------------------
master.aliveWorkers
4/1/20 12:13:40 PM =============================================================

-- Gauges ----------------------------------------------------------------------
master.aliveWorkers
4/1/20 12:33:40 PM =============================================================

-- Gauges ----------------------------------------------------------------------
master.aliveWorkers
4/1/20 12:53:40 PM =============================================================

I'm not sure what those tests are from. @dongjoon-hyun would you happen to know?

dongjoon-hyun · 2020-04-01T20:44:51Z

Ur, is this a consistent failure on this PR? Let me check. I saw the same failure time to time before, but not consistently.

python/pyspark/rdd.py

dongjoon-hyun · 2020-04-01T20:49:13Z

Retest this please.

core/src/main/scala/org/apache/spark/resource/TaskResourceRequests.scala

core/src/main/scala/org/apache/spark/resource/ExecutorResourceRequests.scala

SparkQA · 2020-04-02T03:36:49Z

Test build #120692 has finished for PR 28085 at commit 6a90fbe.

This patch fails from timeout after a configured wait of 400m.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2020-04-02T03:44:23Z

Test build #120693 has finished for PR 28085 at commit 2d754f7.

This patch fails from timeout after a configured wait of 400m.
This patch merges cleanly.
This patch adds no public classes.

dongjoon-hyun · 2020-04-02T18:30:16Z

The running Jenkins job still seems to fail.

-- Gauges ----------------------------------------------------------------------
master.aliveWorkers
4/2/20 9:15:05 AM ==============================================================

dongjoon-hyun · 2020-04-02T18:30:53Z

For now, I have no clue~, @tgravescs .

tgravescs · 2020-04-02T18:40:19Z

do you know which test is outputting that?

SparkQA · 2020-04-02T20:14:17Z

Test build #120723 has finished for PR 28085 at commit 1071f40.

This patch fails from timeout after a configured wait of 400m.
This patch merges cleanly.
This patch adds no public classes.

hit some issues with switching between objects created before SparkContext and ones after. Easier to understand this way. Add tests

tgravescs · 2020-04-16T20:50:54Z

@HyukjinKwon I believe I addressed all your comments. I made it so the python classes can be called before SparkContext creation with the exception of the ResourceProfile.id function.

SparkQA · 2020-04-16T21:46:58Z

Test build #121377 has finished for PR 28085 at commit a6e9ac2.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds no public classes.

tgravescs · 2020-04-16T21:55:59Z

test this please

SparkQA · 2020-04-17T01:10:24Z

Test build #121383 has finished for PR 28085 at commit a6e9ac2.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

HyukjinKwon · 2020-04-20T00:57:51Z

That's nice, thanks @tgravescs for addressing my comments.

python/pyspark/resource/tests/__init__.py

python/pyspark/resource/resourceprofile.py

HyukjinKwon · 2020-04-20T01:04:54Z

Looks pretty fine to me. I will do a final review and leave my sign-off within few days.

Ngone51 · 2020-04-20T13:42:39Z

I am basically ok with the change now. But I'm not good at python so it's still depends on @HyukjinKwon .

SparkQA · 2020-04-20T18:23:34Z

Test build #121535 has finished for PR 28085 at commit 528094c.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

tgravescs · 2020-04-20T18:35:37Z

looks like resource pyspark test ran now. Let me know if I missed anything with that.

Starting test(pypy): pyspark.resource.tests.test_resources
Finished test(pypy): pyspark.profiler (8s)
Starting test(pypy): pyspark.serializers
Finished test(pypy): pyspark.resource.tests.test_resources (6s)

HyukjinKwon · 2020-04-22T03:33:46Z

retest this please

HyukjinKwon

LGTM. several nits.

core/src/main/scala/org/apache/spark/scheduler/DAGScheduler.scala

python/pyspark/resource/executorrequests.py

python/pyspark/rdd.py

python/pyspark/resource/executorrequests.py

SparkQA · 2020-04-22T06:05:20Z

Test build #121604 has finished for PR 28085 at commit 528094c.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2020-04-22T16:47:11Z

Test build #121624 has finished for PR 28085 at commit 89be02e.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2020-04-22T18:05:14Z

Test build #121627 has finished for PR 28085 at commit 354fb0c.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

HyukjinKwon · 2020-04-23T01:20:12Z

Merged to master.

Thank you @tgravescs.

The DecommissionWorkerSuite started becoming flaky and it revealed a real regression. Recent PR's (apache#28085 and apache#29211) neccessitate a small reworking of the decommissioning logic. Before getting into that, let me describe the intended behavior of decommissioning: If a fetch failure happens where the source executor was decommissioned, we want to treat that as an eager signal to clear all shuffle state associated with that executor. In addition if we know that the host was decommissioned, we want to forget about all map statuses from all other executors on that decommissioned host. This is what the test "decommission workers ensure that fetch failures lead to rerun" is trying to test. This invariant is important to ensure that decommissioning a host does not lead to multiple fetch failures that might fail the job. - Per apache#29211, the executors now eagerly exit on decommissioning and thus the executor is lost before the fetch failure even happens. (I tested this by waiting some seconds before triggering the fetch failure). When an executor is lost, we forget its decommissioning information. The fix is to keep the decommissioning information around for some time after removal with some extra logic to finally purge it after a timeout. - Per apache#28085, when the executor is lost, it forgets the shuffle state about just that executor and increments the shuffleFileLostEpoch. This incrementing precludes the clearing of state of the entire host when the fetch failure happens. I elected to only change this codepath for the special case of decommissioning, without any other side effects. This whole version keeping stuff is complex and it has effectively not been semantically changed since 2013! The fix here is also simple: Ignore the shuffleFileLostEpoch when the shuffle status is being cleared due to a fetch failure resulting from host decommission. These two fixes are local to decommissioning only and don't change other behavior. I also added some more tests to TaskSchedulerImpl to ensure that the decommissioning information is indeed purged after a timeout.

The DecommissionWorkerSuite started becoming flaky and it revealed a real regression. Recent PR's (apache#28085 and apache#29211) neccessitate a small reworking of the decommissioning logic. Before getting into that, let me describe the intended behavior of decommissioning: If a fetch failure happens where the source executor was decommissioned, we want to treat that as an eager signal to clear all shuffle state associated with that executor. In addition if we know that the host was decommissioned, we want to forget about all map statuses from all other executors on that decommissioned host. This is what the test "decommission workers ensure that fetch failures lead to rerun" is trying to test. This invariant is important to ensure that decommissioning a host does not lead to multiple fetch failures that might fail the job. - Per apache#29211, the executors now eagerly exit on decommissioning and thus the executor is lost before the fetch failure even happens. (I tested this by waiting some seconds before triggering the fetch failure). When an executor is lost, we forget its decommissioning information. The fix is to keep the decommissioning information around for some time after removal with some extra logic to finally purge it after a timeout. - Per apache#28085, when the executor is lost, it forgets the shuffle state about just that executor and increments the shuffleFileLostEpoch. This incrementing precludes the clearing of state of the entire host when the fetch failure happens. This PR elects to only change this codepath for the special case of decommissioning, without any other side effects. This whole version keeping stuff is complex and it has effectively not been semantically changed since 2013! The fix here is also simple: Ignore the shuffleFileLostEpoch when the shuffle status is being cleared due to a fetch failure resulting from host decommission. These two fixes are local to decommissioning only and don't change other behavior. Also added some more tests to TaskSchedulerImpl to ensure that the decommissioning information is indeed purged after a timeout. Also hardened the test DecommissionWorkerSuite to make it wait for successful job completion.

tgravescs added 8 commits March 30, 2020 14:31

Stage level scheduling python api support

acdf6e8

revert pom changes

24c1a96

Fix log messages

5647535

Try changing way we pass pyspark memory

af69e4b

Change to use local property to pass pyspark memory

4a7f39a

add missing api to get java map

27e1a10

Add java api test

af602b6

cleanup

2b515c8

fix indentation

a052427

fix newline around markup in python

6a90fbe

dongjoon-hyun added the PYSPARK label Apr 1, 2020

dongjoon-hyun reviewed Apr 1, 2020

View reviewed changes

python/pyspark/rdd.py Outdated Show resolved Hide resolved

dongjoon-hyun reviewed Apr 1, 2020

View reviewed changes

core/src/main/scala/org/apache/spark/resource/TaskResourceRequests.scala Outdated Show resolved Hide resolved

dongjoon-hyun reviewed Apr 1, 2020

View reviewed changes

core/src/main/scala/org/apache/spark/resource/ExecutorResourceRequests.scala Outdated Show resolved Hide resolved

Update the version added for rdd api's

2d754f7

make java return values immutable

1071f40

Change to make python versions do same thing as the scala versions as

a6e9ac2

hit some issues with switching between objects created before SparkContext and ones after. Easier to understand this way. Add tests

HyukjinKwon reviewed Apr 20, 2020

View reviewed changes

python/pyspark/resource/tests/__init__.py Show resolved Hide resolved

HyukjinKwon reviewed Apr 20, 2020

View reviewed changes

python/pyspark/resource/resourceprofile.py Show resolved Hide resolved

add pyspark resource module to testing module

528094c

probot-autolabeler bot added the BUILD label Apr 20, 2020

HyukjinKwon approved these changes Apr 22, 2020

View reviewed changes

tgravescs added 2 commits April 22, 2020 08:47

Update names of function/variable

89be02e

Other variable name changes

354fb0c

HyukjinKwon closed this in 95aec09 Apr 23, 2020

zero323 mentioned this pull request May 18, 2020

[SPARK-29641] Stage Level Sched: Add python api's and tests zero323/pyspark-stubs#404

Closed

agrawaldevesh mentioned this pull request Aug 13, 2020

[SPARK-32613][CORE] Fix regressions in DecommissionWorkerSuite #29422

Closed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[SPARK-29641][PYTHON][CORE] Stage Level Sched: Add python api's and tests #28085

[SPARK-29641][PYTHON][CORE] Stage Level Sched: Add python api's and tests #28085

tgravescs commented Mar 31, 2020

SparkQA commented Mar 31, 2020

SparkQA commented Mar 31, 2020

SparkQA commented Apr 1, 2020

tgravescs commented Apr 1, 2020

SparkQA commented Apr 1, 2020

tgravescs commented Apr 1, 2020

dongjoon-hyun commented Apr 1, 2020 •

edited

dongjoon-hyun commented Apr 1, 2020

SparkQA commented Apr 2, 2020

SparkQA commented Apr 2, 2020

dongjoon-hyun commented Apr 2, 2020

dongjoon-hyun commented Apr 2, 2020

tgravescs commented Apr 2, 2020

SparkQA commented Apr 2, 2020

tgravescs commented Apr 16, 2020

SparkQA commented Apr 16, 2020

tgravescs commented Apr 16, 2020

SparkQA commented Apr 17, 2020

HyukjinKwon commented Apr 20, 2020

HyukjinKwon commented Apr 20, 2020

Ngone51 commented Apr 20, 2020 •

edited

SparkQA commented Apr 20, 2020

tgravescs commented Apr 20, 2020

HyukjinKwon commented Apr 22, 2020

HyukjinKwon left a comment

SparkQA commented Apr 22, 2020

SparkQA commented Apr 22, 2020

SparkQA commented Apr 22, 2020

HyukjinKwon commented Apr 23, 2020

[SPARK-29641][PYTHON][CORE] Stage Level Sched: Add python api's and tests #28085

[SPARK-29641][PYTHON][CORE] Stage Level Sched: Add python api's and tests #28085

Conversation

tgravescs commented Mar 31, 2020

What changes were proposed in this pull request?

Why are the changes needed?

Does this PR introduce any user-facing change?

How was this patch tested?

SparkQA commented Mar 31, 2020

SparkQA commented Mar 31, 2020

SparkQA commented Apr 1, 2020

tgravescs commented Apr 1, 2020

SparkQA commented Apr 1, 2020

tgravescs commented Apr 1, 2020

dongjoon-hyun commented Apr 1, 2020 • edited

dongjoon-hyun commented Apr 1, 2020

SparkQA commented Apr 2, 2020

SparkQA commented Apr 2, 2020

dongjoon-hyun commented Apr 2, 2020

dongjoon-hyun commented Apr 2, 2020

tgravescs commented Apr 2, 2020

SparkQA commented Apr 2, 2020

tgravescs commented Apr 16, 2020

SparkQA commented Apr 16, 2020

tgravescs commented Apr 16, 2020

SparkQA commented Apr 17, 2020

HyukjinKwon commented Apr 20, 2020

HyukjinKwon commented Apr 20, 2020

Ngone51 commented Apr 20, 2020 • edited

SparkQA commented Apr 20, 2020

tgravescs commented Apr 20, 2020

HyukjinKwon commented Apr 22, 2020

HyukjinKwon left a comment

Choose a reason for hiding this comment

SparkQA commented Apr 22, 2020

SparkQA commented Apr 22, 2020

SparkQA commented Apr 22, 2020

HyukjinKwon commented Apr 23, 2020

dongjoon-hyun commented Apr 1, 2020 •

edited

Ngone51 commented Apr 20, 2020 •

edited